Skip to content
Advertisement

100M–1B

93 results

Text

Fineweb-Edu-Chinese-V2.1

Chinese Fineweb Edu Dataset V2.1 中文 English OpenCSG Community 👾github wechat Twitter 📖Technical Report The Chinese Fineweb Edu…

100M–1B·Apache-2.0·Parquet
MultimodalTabularText

smollm-corpus

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…

100M–1B·ODC-BY·Parquet
Text

MegaMath

MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We…

100M–1B·ODC-BY·Parquet
Text

pile-uncopyrighted

Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The…

100M–1B·Custom / Research-only·JSON
Text

P3

P3

100M–1B·Apache-2.0·Parquet
MultimodalTabularText

AutoMathText-V2

🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset   🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…

100M–1B
MultimodalTabularText

fineweb-edu-fortified

Fineweb-Edu-Fortified The composition of fineweb-edu-fortified, produced by automatically clustering a 500k row sample in Airtrain What is it?…

100M–1B·ODC-BY·Parquet
Text

xP3x

xP3x

100M–1B·Apache-2.0·Parquet
Text

Corpus-200B

WebOrganizer/Corpus-200B Paper Website GitHub This dataset is a pre-processed version of the 1b-1x CommonCrawl pool from DataComps-LM cleaned…

100M–1B
Text

github-code-2025-language-split

📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was…

100M–1B·Custom / Research-only·Parquet
Advertisement