Tesseract OCR

The open-source OCR engine behind most free text recognition tools.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

1 more ways to get Tesseract OCR

Package managers

About Tesseract OCR

Tesseract recognises text in images and PDFs in more than a hundred languages using an LSTM neural network engine, outputting plain text, hOCR, TSV or searchable PDF. It is a library and command line tool that many other applications build on.

Strengths

  • Over a hundred languages, with trainable models
  • Outputs searchable PDFs directly, not just text

Limitations

  • Accuracy depends heavily on image preprocessing you must do yourself
  • No interface of its own; it is a command line engine

Details

Pricing
FreeFree and open source under the Apache licence.
License
Apache-2.0
Developer
The Tesseract contributors
Platforms
Windows, macOS, Linux, Command line
How it runs
Downloadable app
Best suited for
The open-source OCR engine behind most free text recognition tools
Categories
PDF tools
Last verified
Added
Provenance
Selected from the TechWalrus Resource Hub (PDF & Documents); facts checked against the developer's own pages, 3 sources on file.

Report a wrong fact or a dead link on this listing