Datasets
911 results
smollm-corpus
SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…
Crowd Recital Yiddish – Source Dataset
Crowd Recital Yiddish - Source Dataset
Goodreads-Books
Dataset Card for "BrightData/Goodreads-Books" Dataset Summary Explore a collection of millions of books with the Goodreads dataset, comprising…
L2D
TL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot…
common_corpus
Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…
shofo-tiktok-general-small
Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive…
AutoMathText-V2
🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset 🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…
agibot_alpha_v30
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…
RealCam-Vid
RealCam-Vid Dataset News 25/04/08: We provide torch dataset demo code for example usage of our RealCam-Vid. 25/03/26: Release…
fineweb-edu-fortified
Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it?…