Trafilatura
Trafilatura crawls web pages and extracts their main text, metadata, and other content.
By Adrien Barbaresi and the Trafilatura contributors
Compare Trafilatura with another product
Read the Trafilatura documentation
View the Trafilatura source code
Alternatives
See all alternativesAbout Trafilatura
Trafilatura is a Python package and command-line tool for discovering web content, downloading pages, and extracting text and metadata. It can process live URLs, downloaded HTML, or parsed HTML, and export results as text, Markdown, CSV, JSON, HTML, XML, or XML-TEI.
It suits researchers and developers building web corpora or processing pages for analysis. The project is published under the Apache 2.0 license.
Strengths
- Extracts main text, metadata, comments, and other page elements
- Supports live pages and previously downloaded HTML
- Exports to text, Markdown, CSV, JSON, HTML, and XML formats
Limitations
- Requires Python and its package dependencies
Details
- Pricing
- FreeFree and open source under the Apache 2.0 license.
- License
- Apache-2.0
- Developer
- Adrien Barbaresi and the Trafilatura contributors
- Platforms
- Windows, macOS, Linux, Command line
- How it runs
- Downloadable app
- Works offline
- Yes
- Best suited for
- Researchers and developers collecting clean web text and metadata
- Categories
- CLI tools, Developer tools
- Last verified
- Added
- Provenance
- Selected from the Homebrew formula