ODC-BY
47 results
MegaMath-Web-Pro-Max
OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample…
molmobot-data
MolmoBot-data Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms:…
finemath
📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+)…
FineWeb-Edu 100BT (Shuffled)
FineWeb-Edu 100BT (Shuffled)
smollm-corpus
SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…
fineweb-edu-translated
Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…
fineweb-edu-fortified
Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it?…
S2ORC Full (Semantic Scholar Open Research Corpus)
S2ORC Full (Semantic Scholar Open Research Corpus)
FineWeb Tokenized (AnisoleAI)
FineWeb Tokenized (AnisoleAI)