Custom / Research-only
363 results
MIT_environmental_impulse_responses
MIT Environmental Impulse Response Dataset The audio recordings in this dataset are originally created by the Computational Audition…
Crowd Whatsapp Yiddish – Source Dataset
Crowd Whatsapp Yiddish - Source Dataset
Knesset Plenum Source Dataset
Knesset Plenum Source Dataset
NatureLM-audio-training
NatureLM-audio-training
WideSearch
WideSearch: Benchmarking Agentic Broad Info-Seeking Dataset Summary WideSearch is a benchmark designed to evaluate the capabilities of Large…
stsbenchmark-sts
STSBenchmark An MTEB dataset Massive Text Embedding Benchmark Semantic Textual Similarity Benchmark (STSbenchmark) dataset. Task category t2t Domains…
fev_datasets
Forecast evaluation datasets This repository contains time series datasets that can be used for evaluation of univariate &…
X-Voice-Dataset-Train
X-Voice Training Dataset Overview The X-Voice training dataset is a large-scale multilingual speech corpus curated for high-performance speech…
Russian Podcasts (unlabeled)
Russian Podcasts (unlabeled)
Danbooruwildcards
This is a set of wildcards for danbooru tags. Artist:Prompts for random artist styles, covering approximately 0.6M different…
Stanford Sentiment Treebank v2
Stanford Sentiment Treebank v2
EuroSpeech
EuroSpeech Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary speech…
webdev-arena-preference-10k
WebDev Arena Preference Dataset This dataset contains 10K real-world Webdev Arena battle with 10 state-of-the-art LLMs. More details…
Crowd Recital Yiddish – Source Dataset
Crowd Recital Yiddish - Source Dataset