Skip to content
Advertisement
MultimodalTabularText

📄 FinePDFs

📄 FinePDFs

Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. 📄 FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page:

Source: Hugging Face Hub (HuggingFaceFW/finepdfs). Metadata imported from the dataset’s Hub tags.

Advertisement