Skip to content
Advertisement
MultimodalTabularText

πŸ“„ FinePDFs

πŸ“„ FinePDFs

Liberating 3T of the finest tokens from PDFs What is this? As we run out of web pages to process, the natural question has always been: what to do next? Only a few knew about a data source that everyone avoided for ages, due to its incredible extraction cost and complexity: PDFs. πŸ“„ FinePDFs is exactly that. It is the largest publicly available corpus sourced exclusively from PDFs, containing about 3 trillion tokens across 475 million documents in 1733 languages. Compared to… See the full description on the dataset page:

Source: Hugging Face Hub (HuggingFaceFW/finepdfs). Metadata imported from the dataset’s Hub tags.

Advertisement