trafilatura Open source
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON,
Collection & ScrapingPython & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Open-source data tool (Apache-2.0 license, 6,328★). Source: GitHub (adbar/trafilatura).
Similar tools
Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability manageme
MultimodalA powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
TextFrom the blog
Avoiding Data Leakage in Synthetic Data Projects
Synthetic data offers vast opportunities for machine learning model training without real-world constraints. Yet, beneath this promise lies the risk of…
Read more →Practical Guide to Scaling Synthetic Data Solutions
Imagine handling a growing machine learning project with increasing demands for training data. Traditional data methods can't keep up, making synthetic…
Read more →Mastering Synthetic Data Generation: From Concepts to Deployment
Synthetic data is more than a trend; it's a crucial tool for building strong AI models, especially when real-world data is scarce, sensitive, or…
Read more →