MinerU

A document parser that converts PDFs, Office files and scans into Markdown and JSON.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About MinerU

MinerU converts complex documents into structured Markdown and JSON suitable for language models and search. It handles PDFs, scanned images, Word, PowerPoint and Excel files, OpenDocument formats, EPUB, OFD, HTML and CSV, and deals with tables, formulas and OCR.

Version 4.0 offers four parsing tiers from fast previews to high-quality layout analysis, a local document library with caching, search and page-level citation locators, and a choice of model backends including ONNX, Torch, llama.cpp, vLLM and LMDeploy. It can be used through a Python SDK, an API or batch conversion.

Strengths

  • Parses tables, formulas and scanned pages with OCR
  • Accepts PDF, Office, OpenDocument, EPUB, HTML and more
  • Several local model backends, from ONNX to vLLM
  • Very actively developed

Limitations

  • Setup and model downloads suit technical users
  • Higher-quality parsing tiers need more compute
  • License is not a standard OSI license as declared on GitHub

Details

Pricing
FreeFree to download and run; the repository does not declare a standard license.
License
Proprietary
Developer
OpenDataLab
Platforms
Self-hosted, Command line
How it runs
Downloadable app, Self-hosted
Account
Not required
Works offline
Yes
Best suited for
Turning document collections into clean Markdown for AI and search pipelines
Last verified
Added
Provenance
Facts checked against the developer's own pages and store listings, 1 sources on file.

Alternatives to MinerU

Compare all

Software that can replace MinerU for an important use case, and what changes if you switch.

  • Marker

    An open-source tool that converts PDFs and other documents to Markdown, JSON and HTML.

    Marker converts PDFs and other documents to Markdown, JSON and HTML under an Apache-2.0 code licence, with model weights under a separate licence.

  • olmOCR

    Open-source toolkit that uses a vision language model to turn PDFs and scans into Markdown.

    olmOCR turns PDFs and scans into Markdown using an open vision language model and keeps reading order, but requires a capable GPU and Linux.

  • MarkItDown

    Python command-line tool that converts PDFs, Office files and other documents into Markdown.

    MarkItDown is a lighter MIT Python tool converting PDFs and Office files to Markdown, aimed at LLM input rather than high-fidelity parsing of formulas and scans.

  • PaddleOCR

    OCR toolkit that turns PDFs and images into structured data.

    PaddleOCR is Apache-2.0 licensed and outputs structured JSON and Markdown from layout, table and formula parsing, relying on PaddlePaddle.

MinerU as an alternative

Listings that name MinerU as an alternative.

  • Docling

    Open source library that converts documents into AI-ready structured formats.

    MinerU converts PDFs, Office files and scans into Markdown and JSON with OCR and several local backends, but its licence is not a standard OSI licence.

  • EasyOCR

    Python OCR library covering more than eighty languages out of the box.

    MinerU parses PDFs, Office files and scans into Markdown and JSON with OCR, offering several local model backends, though its license is not a standard OSI license.

  • pdf2htmlEX

    A command-line converter that turns PDF files into HTML while keeping text and layout.

    MinerU parses PDFs and scans into Markdown and JSON for pipelines, with OCR support, rather than producing layout-faithful HTML pages.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing