Skip to content
Advertisement

ODC-BY

47 results

MultimodalTabularText

MegaMath-Web-Pro-Max

OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling The Curation of MegaMath-Web-Pro-Max Step 1: Uniformly and randomly sample…

10M–100M·ODC-BY·Parquet
MultimodalTabularText

molmobot-data

MolmoBot-data Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms:…

100K–1M·ODC-BY·Parquet
MultimodalTabularText

finemath

📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+)…

10M–100M·ODC-BY·Parquet
MultimodalTabularText

smollm-corpus

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…

100M–1B·ODC-BY·Parquet
Text

fineweb-edu-translated

Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…

>1B·ODC-BY·Parquet
Text

MegaMath

MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We…

100M–1B·ODC-BY·Parquet
MultimodalTabularText

fineweb-edu-fortified

Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it?…

100M–1B·ODC-BY·Parquet
Advertisement