Skip to content
Advertisement

Datasets

3,216 results

Text

SWE-bench_Pro

Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…

<1K·Parquet
Text

TriviaQA

TriviaQA

100K–1M·Custom / Research-only·Parquet
Text

P3

P3

100M–1B·Apache-2.0·Parquet
MultimodalTabularText

common_corpus

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…

10K–100K·Parquet
Text

pile-uncopyrighted

Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The…

100M–1B·Custom / Research-only·JSON
Text

wikitext_document_level

Wikitext Document Level This is a modified version of that returns Wiki pages instead of Wiki text line-by-line.…

10K–100K·CC-BY-SA·Parquet
Text

OpenThoughts-114k

!NOTE We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-114k Open synthetic reasoning dataset with…

100K–1M·Apache-2.0·Parquet
MultimodalTabularText

AutoMathText-V2

🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset   🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…

100M–1B
Text

SWE-bench_Verified

Dataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been…

<1K·Parquet
MultimodalTabularText

agibot_alpha_v30

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…

10M–100M·Apache-2.0·Parquet
Text

TinyStories

Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in…

1M–10M·Other·Parquet
Text

Open-Sora-Plan-v1.1.0

Annotation We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match…

100K–1M·MIT·WebDataset
Advertisement