Skip to content
Advertisement

Text

2,175 results

MultimodalTabularText

smollm-corpus

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…

100M–1B·ODC-BY·Parquet
Text

webdev-arena-preference-10k

WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details…

10K–100K·Custom / Research-only·JSON
Text

MegaMath

MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We…

100M–1B·ODC-BY·Parquet
Text

SWE-bench

Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects…

10K–100K·Parquet
Text

fineweb-edu-translated

Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…

>1B·ODC-BY·Parquet
Text

recipes

🦛 Chonkie Recipes 🍳 Chonkie loves to cook up a storm in the kitchen This repository contains all…

<1K·Apache-2.0
Text

Vchitect_T2V_DataVerse

Vchitect-T2V-Dataverse Vchitect Team1 1Shanghai Artificial Intelligence Laboratory Paper Project Page Data Overview The Vchitect-T2V-Dataverse is the…

1M–10M·Apache-2.0·WebDataset
Text

LongBench-v2

LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: 💻 Github Repo: 📚…

<1K·Apache-2.0·JSON
Text

WMT T2T

WMT T2T

1M–10M·Custom / Research-only·Parquet
AudioMultimodalText

yodas2_sidon

YODAS2-Sidon Overview This dataset is a cleansed version of YODAS-2 with Sidon speech restoration mode for Speech Synthesis…

1M–10M·CC-BY·WebDataset
Text

BoolQ

BoolQ

10K–100K·CC-BY-SA·Parquet
Advertisement