Trafilatura

Trafilatura crawls web pages and extracts their main text, metadata, and other content.

By Adrien Barbaresi and the Trafilatura contributors

Compare Trafilatura with another product

Read the Trafilatura documentation

View the Trafilatura source code

Alternatives

See all alternatives

About Trafilatura

Trafilatura is a Python package and command-line tool for discovering web content, downloading pages, and extracting text and metadata. It can process live URLs, downloaded HTML, or parsed HTML, and export results as text, Markdown, CSV, JSON, HTML, XML, or XML-TEI.

It suits researchers and developers building web corpora or processing pages for analysis. The project is published under the Apache 2.0 license.

Strengths

  • Extracts main text, metadata, comments, and other page elements
  • Supports live pages and previously downloaded HTML
  • Exports to text, Markdown, CSV, JSON, HTML, and XML formats

Limitations

  • Requires Python and its package dependencies

Details

Pricing
FreeFree and open source under the Apache 2.0 license.
License
Apache-2.0
Developer
Adrien Barbaresi and the Trafilatura contributors
Platforms
Windows, macOS, Linux, Command line
How it runs
Downloadable app
Works offline
Yes
Best suited for
Researchers and developers collecting clean web text and metadata
Last verified
Added
Provenance
Selected from the Homebrew formula

Report a wrong fact or a dead link on this listing