Skip to content
Advertisement

Custom / Research-only

363 results

Audio

MIT_environmental_impulse_responses

MIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition…

<1K·Custom / Research-only·Audio (folder)
Text

WideSearch

WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large…

<1K·Custom / Research-only·JSON
Text

stsbenchmark-sts

STSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset. Task category t2t Domains…

1K–10K·Custom / Research-only·JSON
AudioMultimodalText

X-Voice-Dataset-Train

X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech…

10M–100M·Custom / Research-only·WebDataset
Text

TweetEval

TweetEval

100K–1M·Custom / Research-only·Parquet
Text

Danbooruwildcards

This is a set of wildcards for danbooru tags. Artist:Prompts for random artist styles, covering approximately 0.6M different…

10M–100M·Custom / Research-only·Text (raw)
Text

sts12-sts

STS12 An MTEB dataset Massive Text Embedding Benchmark SemEval-2012 Task 6. Task category t2t Domains Encyclopaedic, News, Written…

1K–10K·Custom / Research-only·JSON
Text

Emotion

Emotion

100K–1M·Custom / Research-only·Parquet
AudioMultimodalText

EuroSpeech

EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…

10M–100M·Custom / Research-only·Parquet
Text

webdev-arena-preference-10k

WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details…

10K–100K·Custom / Research-only·JSON
Text

WMT T2T

WMT T2T

1M–10M·Custom / Research-only·Parquet
Advertisement