pdf2htmlEX

A command-line converter that turns PDF files into HTML while keeping text and layout.

These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.

About pdf2htmlEX

pdf2htmlEX converts PDF documents into HTML pages that look like the original, with text kept as real, selectable text rather than images. This fork brings together fixes from other forks to keep the project active and maintained.

It adds more accurate handling of text that is fully or partially covered by other content, some support for transparent text, and DPI limits that stop output graphics from growing too large. Options control font sizing, zoom and the resolution used for rasterized text.

Strengths

  • Keeps text selectable and searchable in the HTML output
  • Close visual match to the original layout
  • Handles text that is hidden or partly hidden behind other content

Limitations

  • Command line only, with no graphical interface
  • No packaged releases on GitHub
  • Getting exact results can mean tuning several options

Details

Pricing
FreeFree and open source.
License
GPL-3.0
Developer
The pdf2htmlEX contributors
Platforms
Linux, Command line
How it runs
Downloadable app
Account
Not required
Works offline
Yes
Best suited for
Publishing PDFs on the web as searchable HTML pages
Categories
PDF tools, CLI tools
Last verified
Added
Provenance
Facts checked against the developer's own pages and store listings, 2 sources on file.

Alternatives to pdf2htmlEX

Compare all

Software that can replace pdf2htmlEX for an important use case, and what changes if you switch.

  • Xpdf

    A PDF viewer plus command-line tools for extracting text and images and converting PDFs to HTML.

    Xpdf includes command-line tools for converting PDFs to HTML and extracting text and images, plus a lightweight viewer, though not every component is open source.

  • Marker

    An open-source tool that converts PDFs and other documents to Markdown, JSON and HTML.

    Marker converts PDFs to HTML, Markdown and JSON with tables and equations, but outputs clean structure rather than a visual copy of the layout.

  • MinerU

    A document parser that converts PDFs, Office files and scans into Markdown and JSON.

    MinerU parses PDFs and scans into Markdown and JSON for pipelines, with OCR support, rather than producing layout-faithful HTML pages.

Similar software

Related functionality, not necessarily a direct replacement.

Report a wrong fact or a dead link on this listing