Skip to content
Advertisement

Parquet

1,501 results

Text

Emotion

Emotion

100K–1M·Custom / Research-only·Parquet
Text

pile-10k

The first 10K elements of The Pile, useful for debugging models trained on it. See the HuggingFace page…

10K–100K·Other·Parquet
Text

OpenR1-Math-220k

OpenR1-Math-220k Dataset description OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with…

100K–1M·Apache-2.0·Parquet
Text

SWE-Gym

SWE-Gym contains 2438 instances sourced from 11 Python repos, following SWE-Bench data collection procedure. Get started at project…

1K–10K·MIT·Parquet
AudioMultimodalText

EuroSpeech

EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…

10M–100M·Custom / Research-only·Parquet
Text

RMISC

RMISC: A Large-scale Real-world Multivariate Corpus for Time Series Foundation Models This dataset card describes the main branch…

>1B·MIT·Parquet
Text

nemotron-cc-translated

Helsinki-NLP/nemotron-cc-translated nemotron-cc-tanslated is a collection of automatically translated documents from nemotron-cc taken out of the…

>1B·CC0·Parquet
AudioMultimodalText

covost2

This is a partial copy of CoVoST2 dataset. The main difference is that the audio data is included…

1M–10M·Parquet
Advertisement