OCRmyPDF
Adds a searchable text layer to a scanned PDF without disturbing the page images, producing a PDF/A you can search and copy from.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
1 more ways to get OCRmyPDF
Package managers
- Homebrew
brew install ocrmypdf
About OCRmyPDF
A scanned PDF is a stack of pictures; you cannot search it or copy text out of it. OCRmyPDF runs Tesseract over the pages and places the recognised text invisibly underneath the images, so the document looks identical but is now searchable and selectable.
It is careful about the things that matter: it keeps the exact resolution of the original embedded images, inserts the OCR layer losslessly where possible so nothing else in the file is disturbed, and validates both input and output. It also optimises images, often producing a smaller file than it started with, and can deskew or clean pages before recognition.
It uses all available CPU cores, handles files with thousands of pages, supports over 100 languages through Tesseract, and does all of it locally.
Strengths
- Places OCR text accurately below the image, so copy and paste lands in the right place
- Keeps the original image resolution and inserts the text layer losslessly where possible
- Often produces a smaller file than the input by optimising images
- Over 100 languages through Tesseract, and it all runs locally
Limitations
- Command line only
- Accuracy depends on the scan; poor originals produce poor text
- Tesseract and Ghostscript are separate dependencies to install
- Processing large batches is CPU-intensive and slow
Details
- Pricing
- FreeFree and open source under the Mozilla Public Licence 2.0.
- License
- MPL-2.0
- Developer
- OCRmyPDF
- Platforms
- Windows, macOS, Linux, Command line
- How it runs
- Downloadable app
- Account
- Not required
- Works offline
- Yes
- Best suited for
- Making a pile of scanned PDFs searchable without changing how they look
- Categories
- Office tools
- Last verified
- Added
- Provenance
- Selected from the TechWalrus downloads catalog (Office); facts checked against the developer's own pages, 3 sources on file.
Alternatives to OCRmyPDF
Compare allSoftware that can replace OCRmyPDF for an important use case, and what changes if you switch.
gImageReader
Recognize text from scans and PDFs through a desktop interface.
gImageReader is a free GPL-3.0 desktop interface for recognizing text from scans and PDFs on Windows and Linux, with editable recognition areas instead of command-line use.
Similar software
Related functionality, not necessarily a direct replacement.
Paperless-ngx
Turns scanned paper into a searchable archive: OCR, automatic tagging, full-text search and a web interface you host yourself.
Pot
Cross-platform translation and OCR tool triggered by selecting text.
Teedy
Lightweight self-hosted document management server.
Mayan EDMS
Mature open-source document management system with workflows.