Skip to content
Advertisement

100K–1M

595 results

AudioMultimodalText

MRSAudio

MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…

100K–1M·CC-BY·CSV
MultimodalTextVideo

FLARE

FLARE: Full-Modality Long-Video Audiovisual Retrieval Benchmark with User-Simulated Queries 🤗 About This Benchmark This repository hosts the full…

100K–1M·CC-BY·JSON
Text

hh-rlhf

Dataset Card for HH-RLHF Dataset Summary This repository provides access to two different kinds of data: Human preference…

100K–1M·MIT·JSON
Text

nfcorpus

NFCorpus An MTEB dataset Massive Text Embedding Benchmark NFCorpus: A Full-Text Learning to Rank Dataset for Medical Information…

100K–1M·JSON
Text

TweetEval

TweetEval

100K–1M·Custom / Research-only·Parquet
Text

Emotion

Emotion

100K–1M·Custom / Research-only·Parquet
Text

OpenR1-Math-220k

OpenR1-Math-220k Dataset description OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with…

100K–1M·Apache-2.0·Parquet
Text

MetaMathQA

View the project page: see our paper at Note All MetaMathQA data are augmented from the training sets…

100K–1M·MIT·JSON
Advertisement