
trafilatura 2.0.0
0
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, o
Contents
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML.
Stars: 5310, Watchers: 5310, Forks: 343, Open Issues: 97The adbar/trafilatura repo was created 6 years ago and the last code push was 5 months ago.
The project is extremely popular with a mindblowing 5310 github stars!
How to Install trafilatura
You can install trafilatura using pip
pip install trafilatura
or add it to a project with poetry
poetry add trafilatura
Package Details
- Author
- None
- License
- Apache 2.0
- Homepage
- None
- PyPi:
- https://pypi.org/project/trafilatura/
- GitHub Repo:
- https://github.com/adbar/trafilatura
Classifiers
- Internet/WWW/HTTP
- Scientific/Engineering/Information Analysis
- Security
- Text Editors/Text Processing
- Text Processing/Linguistic
- Text Processing/Markup/HTML
- Text Processing/Markup/Markdown
- Text Processing/Markup/XML
- Utilities
Related Packages
Errors
A list of common trafilatura errors.
Code Examples
Here are some trafilatura code examples and snippets.
GitHub Issues
The trafilatura package has 97 open issues on GitHub
- 🚨 URGENT: Your TogetherAI key is having a public meltdown! ðŸŽ
- favor_precision and include_links Enabled Together Cause Over-Aggressive Noise Removal and Main Content Misidentification
- Feature request: option to keep <span> tags (or disable strip_tags for span) in HTML/Markdown output
- Table tags incorrect in HTML formatted output
- Enhancements: Regex Extraction, Advanced Flags, and CLI Improvements
- Duplicate paragraph extraction when a long sibling paragraph is present
- Add Kurdish language as it is available in the Justext
- fix: table handling and img tags inside tables
pythonfix







