Skip to content
Advertisement

Datasets

911 results

MultimodalTabularText

smollm-corpus

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…

100M–1B·ODC-BY·Parquet
MultimodalTabularText

Goodreads-Books

Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…

1M–10M·Custom / Research-only·CSV
MultimodalTabularText

common_corpus

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…

10K–100K·Parquet
MultimodalTabularText

AutoMathText-V2

🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset   🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…

100M–1B
MultimodalTabularText

agibot_alpha_v30

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…

10M–100M·Apache-2.0·Parquet
MultimodalTabularText

RealCam-Vid

RealCam-Vid Dataset News 25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid. 25/03/26: Release…

100K–1M·MIT·CSV
MultimodalTabularText

fineweb-edu-fortified

Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it?…

100M–1B·ODC-BY·Parquet
Advertisement