PaddleOCR

OCR toolkit that turns PDFs and images into structured data.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About PaddleOCR

PaddleOCR detects and recognises text in more than a hundred languages and goes further by parsing layout, tables and formulas into structured JSON or Markdown suitable for feeding language models. It offers both lightweight and high-accuracy model sets.

Strengths

  • Parses layout, tables and formulas, not just lines of text
  • Outputs JSON and Markdown ready for downstream processing

Limitations

  • Built on the PaddlePaddle framework, which is a large dependency
  • Much of the documentation and community is in Chinese

Details

Pricing
FreeFree and open source under the Apache licence.
License
Apache-2.0
Developer
PaddlePaddle
Platforms
Windows, macOS, Linux, Command line
How it runs
Downloadable app
Best suited for
OCR toolkit that turns PDFs and images into structured data
Categories
PDF tools
Last verified
Added
Provenance
Selected from the TechWalrus Resource Hub (PDF & Documents); facts checked against the developer's own pages, 3 sources on file.

Report a wrong fact or a dead link on this listing