PaddleOCR
OCR toolkit that turns PDFs and images into structured data.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About PaddleOCR
PaddleOCR detects and recognises text in more than a hundred languages and goes further by parsing layout, tables and formulas into structured JSON or Markdown suitable for feeding language models. It offers both lightweight and high-accuracy model sets.
Strengths
- Parses layout, tables and formulas, not just lines of text
- Outputs JSON and Markdown ready for downstream processing
Limitations
- Built on the PaddlePaddle framework, which is a large dependency
- Much of the documentation and community is in Chinese
Details
- Pricing
- FreeFree and open source under the Apache licence.
- License
- Apache-2.0
- Developer
- PaddlePaddle
- Platforms
- Windows, macOS, Linux, Command line
- How it runs
- Downloadable app
- Best suited for
- OCR toolkit that turns PDFs and images into structured data
- Categories
- PDF tools
- Last verified
- Added
- Provenance
- Selected from the TechWalrus Resource Hub (PDF & Documents); facts checked against the developer's own pages, 3 sources on file.