
ocrmypdf 17.2.0
0
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Contents
OCRmyPDF adds an OCR text layer to scanned PDF files, allowing them to be searched
Stars: 32647, Watchers: 32647, Forks: 2271, Open Issues: 142The ocrmypdf/OCRmyPDF repo was created 12 years ago and the last code push was 1 weeks ago.
The project is extremely popular with a mindblowing 32647 github stars!
How to Install ocrmypdf
You can install ocrmypdf using pip
pip install ocrmypdf
or add it to a project with poetry
poetry add ocrmypdf
Package Details
- Author
- None
- License
- None
- Homepage
- None
- PyPi:
- https://pypi.org/project/ocrmypdf/
- Documentation:
- https://ocrmypdf.readthedocs.io/
- GitHub Repo:
- https://github.com/ocrmypdf/OCRmyPDF
Classifiers
- Scientific/Engineering/Image Recognition
- Text Processing/Indexing
- Text Processing/Linguistic
Related Packages
Errors
A list of common ocrmypdf errors.
Code Examples
Here are some ocrmypdf code examples and snippets.
GitHub Issues
The ocrmypdf package has 142 open issues on GitHub
- Feature: Allow "end" as alias for the last page
- Set work_folder in PdfContext options initialization
- Bug: Ghostscript rasterizing fails on seemingly empty page of document
- Feature: Removal of bad text layer while preserving vector graphics as-is
- Bug: Could not found Ghostscript
- Feature: and support PaddlePaddle OCR in the backend
- Add GVision OCR engine support [NOT MERGEABLE due to unclear licensing]
- Bug: Page not rotating with –rotate-pages (inconsistent results)
- Feature: Integrations with other backends via hOcr (naive implementation of easyOcr backend inside)
pythonfix






