Parquet
1,501 results
smollm-corpus
SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…
fineweb-edu-translated
Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…
evaluation-tables
!CAUTION This dataset will not be updated. It corresponds to the last available public snapshot of the data,…
RottenTomatoes – MR Movie Review Data
RottenTomatoes - MR Movie Review Data
L2D
TL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot…
SWE-bench_Pro
Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…
common_corpus
Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…