pdf2htmlEX
A command-line converter that turns PDF files into HTML while keeping text and layout.
These buttons open the developer's own site, repository or store listing in a new tab. wares.gg does not host downloads.
About pdf2htmlEX
pdf2htmlEX converts PDF documents into HTML pages that look like the original, with text kept as real, selectable text rather than images. This fork brings together fixes from other forks to keep the project active and maintained.
It adds more accurate handling of text that is fully or partially covered by other content, some support for transparent text, and DPI limits that stop output graphics from growing too large. Options control font sizing, zoom and the resolution used for rasterized text.
Strengths
- Keeps text selectable and searchable in the HTML output
- Close visual match to the original layout
- Handles text that is hidden or partly hidden behind other content
Limitations
- Command line only, with no graphical interface
- No packaged releases on GitHub
- Getting exact results can mean tuning several options
Details
- Pricing
- FreeFree and open source.
- License
- GPL-3.0
- Developer
- The pdf2htmlEX contributors
- Platforms
- Linux, Command line
- How it runs
- Downloadable app
- Account
- Not required
- Works offline
- Yes
- Best suited for
- Publishing PDFs on the web as searchable HTML pages
- Last verified
- Added
- Provenance
- Facts checked against the developer's own pages and store listings, 2 sources on file.
Alternatives to pdf2htmlEX
Compare allSoftware that can replace pdf2htmlEX for an important use case, and what changes if you switch.
Xpdf
A PDF viewer plus command-line tools for extracting text and images and converting PDFs to HTML.
Xpdf includes command-line tools for converting PDFs to HTML and extracting text and images, plus a lightweight viewer, though not every component is open source.
Marker
An open-source tool that converts PDFs and other documents to Markdown, JSON and HTML.
Marker converts PDFs to HTML, Markdown and JSON with tables and equations, but outputs clean structure rather than a visual copy of the layout.
MinerU
A document parser that converts PDFs, Office files and scans into Markdown and JSON.
MinerU parses PDFs and scans into Markdown and JSON for pipelines, with OCR support, rather than producing layout-faithful HTML pages.
Similar software
Related functionality, not necessarily a direct replacement.
Poppler
PDF rendering library with pdftotext, pdftoppm and pdfinfo tools.
MarkItDown
Python command-line tool that converts PDFs, Office files and other documents into Markdown.
MuPDF
Compact, very fast PDF and XPS viewer with a command-line toolkit.
mutool (MuPDF)
Command-line companion to MuPDF for PDF extraction, cleaning and merging.
PaddleOCR
OCR toolkit that turns PDFs and images into structured data.
Ghostscript
PostScript and PDF interpreter behind a great deal of other software.