Skip to content
Advertisement

Parquet

1,501 results

Text

aime_2024

Dataset card for AIME 2024 This dataset consists of 30 problems from the 2024 AIME I and AIME…

<1K·Parquet
MultimodalTabularText

smollm-corpus

SmolLM-Corpus This dataset is a curated collection of high-quality educational and synthetic data designed for training small language…

100M–1B·ODC-BY·Parquet
Text

MegaMath

MegaMath: Pushing the Limits of Open Math Copora Megamath is part of TxT360, curated by LLM360 Team. We…

100M–1B·ODC-BY·Parquet
Text

SWE-bench

Dataset Summary SWE-bench is a dataset that tests systems’ ability to solve GitHub issues automatically. The dataset collects…

10K–100K·Parquet
Text

fineweb-edu-translated

Helsinki-NLP/fineweb-edu-translated fineweb-edu-tanslated is a collection of automatically translated documents from fineweb-edu. Translations are…

>1B·ODC-BY·Parquet
Text

WMT T2T

WMT T2T

1M–10M·Custom / Research-only·Parquet
Text

BoolQ

BoolQ

10K–100K·CC-BY-SA·Parquet
Text

C-Eval

C-Eval

10K–100K·CC-BY-NC-SA·Parquet
Text

evaluation-tables

!CAUTION This dataset will not be updated. It corresponds to the last available public snapshot of the data,…

1K–10K·CC-BY-SA·Parquet
Text

SWE-bench_Pro

Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…

<1K·Parquet
Text

TriviaQA

TriviaQA

100K–1M·Custom / Research-only·Parquet
Text

P3

P3

100M–1B·Apache-2.0·Parquet
MultimodalTabularText

common_corpus

Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…

10K–100K·Parquet
Advertisement