Skip to content
Advertisement

Text

2,175 results

Text

pile-10k

The first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page…

10K–100K·Other·Parquet
Text

OpenR1-Math-220k

OpenR1-Math-220k Dataset description OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with…

100K–1M·Apache-2.0·Parquet
Text

MetaMathQA

View the project page: see our paper at Note All MetaMathQA data are augmented from the training sets…

100K–1M·MIT·JSON
Text

SWE-Gym

SWE-Gym contains 2438 instances sourced from 11 Python repos, following SWE-Bench data collection procedure. Get started at project…

1K–10K·MIT·Parquet
Text

aime25

AIME 25 American Invitational Mathematics Examination (AIME) 2025 Citation If you use the AIME25 dataset in your research,…

<1K·Apache-2.0·JSON
Text

databricks-dolly-15k

Summary databricks-dolly-15k is an open source dataset of instruction-following records generated by thousands of Databricks employees in several…

10K–100K·CC-BY-SA·JSON
AudioMultimodalText

EuroSpeech

EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…

10M–100M·Custom / Research-only·Parquet
Text

RMISC

RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models This dataset card describes the main branch…

>1B·MIT·Parquet
Text

nemotron-cc-translated

Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the…

>1B·CC0·Parquet
AudioMultimodalText

covost2

This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included…

1M–10M·Parquet
Text

aime_2024

Dataset card for AIME 2024 This dataset consists of 30 problems from the 2024 AIME I and AIME…

<1K·Parquet
Advertisement