MinerU
A document parser that converts PDFs, Office files and scans into Markdown and JSON.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About MinerU
MinerU converts complex documents into structured Markdown and JSON suitable for language models and search. It handles PDFs, scanned images, Word, PowerPoint and Excel files, OpenDocument formats, EPUB, OFD, HTML and CSV, and deals with tables, formulas and OCR.
Version 4.0 offers four parsing tiers from fast previews to high-quality layout analysis, a local document library with caching, search and page-level citation locators, and a choice of model backends including ONNX, Torch, llama.cpp, vLLM and LMDeploy. It can be used through a Python SDK, an API or batch conversion.
Strengths
- Parses tables, formulas and scanned pages with OCR
- Accepts PDF, Office, OpenDocument, EPUB, HTML and more
- Several local model backends, from ONNX to vLLM
- Very actively developed
Limitations
- Setup and model downloads suit technical users
- Higher-quality parsing tiers need more compute
- License is not a standard OSI license as declared on GitHub
Details
- Pricing
- FreeFree to download and run; the repository does not declare a standard license.
- License
- Proprietary
- Developer
- OpenDataLab
- Platforms
- Self-hosted, Command line
- How it runs
- Downloadable app, Self-hosted
- Account
- Not required
- Works offline
- Yes
- Best suited for
- Turning document collections into clean Markdown for AI and search pipelines
- Categories
- PDF tools, Local AI tools, Developer tools
- Last verified
- Added
- Provenance
- Facts checked against the developer's own pages and store listings, 1 sources on file.
Alternatives to MinerU
Compare allSoftware that can replace MinerU for an important use case, and what changes if you switch.
Marker
An open-source tool that converts PDFs and other documents to Markdown, JSON and HTML.
Marker converts PDFs and other documents to Markdown, JSON and HTML under an Apache-2.0 code licence, with model weights under a separate licence.
olmOCR
Open-source toolkit that uses a vision language model to turn PDFs and scans into Markdown.
olmOCR turns PDFs and scans into Markdown using an open vision language model and keeps reading order, but requires a capable GPU and Linux.
MarkItDown
Python command-line tool that converts PDFs, Office files and other documents into Markdown.
MarkItDown is a lighter MIT Python tool converting PDFs and Office files to Markdown, aimed at LLM input rather than high-fidelity parsing of formulas and scans.
PaddleOCR
OCR toolkit that turns PDFs and images into structured data.
PaddleOCR is Apache-2.0 licensed and outputs structured JSON and Markdown from layout, table and formula parsing, relying on PaddlePaddle.
MinerU as an alternative
Listings that name MinerU as an alternative.
Docling
Open source library that converts documents into AI-ready structured formats.
MinerU converts PDFs, Office files and scans into Markdown and JSON with OCR and several local backends, but its licence is not a standard OSI licence.
EasyOCR
Python OCR library covering more than eighty languages out of the box.
MinerU parses PDFs, Office files and scans into Markdown and JSON with OCR, offering several local model backends, though its license is not a standard OSI license.
pdf2htmlEX
A command-line converter that turns PDF files into HTML while keeping text and layout.
MinerU parses PDFs and scans into Markdown and JSON for pipelines, with OCR support, rather than producing layout-faithful HTML pages.
Similar software
Related functionality, not necessarily a direct replacement.
Tesseract OCR
The open-source OCR engine behind most free text recognition tools.
Poppler
PDF rendering library with pdftotext, pdftoppm and pdfinfo tools.
ChatPDF
An AI tool that summarizes PDFs and answers questions about them, with citations to the source.
Humata
An AI tool for asking questions about your documents, with cited answers and summaries.