Open
2,922 results
music-mood-recs assets (MTG-Jamendo mood/theme subset mirror)
music-mood-recs assets (MTG-Jamendo mood/theme subset mirror)
AppTek Call-Center Dialogues
AppTek Call-Center Dialogues
Tarteel AI – EveryAyah Dataset
Tarteel AI - EveryAyah Dataset
Quranic Recitation Dataset (Word-by-Word)
Quranic Recitation Dataset (Word-by-Word)
AudioMarathon
AudioMarathon AudioMarathon is a long-context audio benchmark for evaluating multimodal LLMs on speech, music, environmental audio, and meetings.…
Afrivoice Swahili Agriculture Subset
Afrivoice Swahili Agriculture Subset
audioset-16khz-wds
AudioSet AudioSet 1 is a large-scale dataset comprising approximately 2 million 10-second YouTube audio clips, categorised into 527…
ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
avspeech-visual-audio
AVSpeech Video + Audio A restructured subset of the AVSpeech dataset with separated media streams and derived identifiers.…
meow-10k
Dataset Card for Meow-10K Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the…
Vietnamese Restaurant Order Speech
Vietnamese Restaurant Order Speech
libritts_r
Dataset Card for LibriTTS-R LibriTTS-R 1 is a sound quality improved version of the LibriTTS corpus ( which…
Bengali Telecom Customer Care Synthetic Speech v2
Bengali Telecom Customer Care Synthetic Speech v2
MMAE
MMAE: A Massive Multitask Audio Editing Benchmark 📖 arXiv 🎬 MMAE Demo Video 🛠️ GitHub Code 🔊 HuggingFace…
AISHELL-3
AISHELL-3 is a large-scale and high-fidelity multi-speaker Mandarin speech corpus published by Beijing Shell Shell Technology Co.,Ltd. It…