SoundingEarth
SoundingEarth SoundingEarth is a geo-referenced soundscape dataset that pairs Google Earth imagery with geotagged environmental audio recordings…
81 results
SoundingEarth SoundingEarth is a geo-referenced soundscape dataset that pairs Google Earth imagery with geotagged environmental audio recordings…
Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…
wuwa-voice-EN wuwa voice EN is a dataset of voice line from Wuthering Waves Attribute Value Language English Total…
Speech DAC Tokens (3 Codebooks)
WorldAudioNaturalConversations Sample Dataset
Risale-i Nur Sesli Külliyat — Risale-i Nur Audio Corpus
Use this dataset in conjuction with: YodasSpeakerPool YodasSpeakerPool is a curated, richly-annotated multi-speaker dataset featuring 7,600 unique…
SibmaBench Data Release & Benchmarking To evaluate your model on SimbaBench across all supported tasks (ASR, TTS, and…
Emolia-Thinking (VoiceNet balanced subset)
Enhanced Audiosnippets Long 2.8M Enhanced version of mitermix/audiosnippets long 2 8M with speech enhancement, emotion annotations, speaker…
C3T: Cross-modal Capabilities Conservation Test Dataset Description C3T (Cross-modal Capabilities Conservation Test) is a benchmark for assessing the…
NaturalVoices Restored (48 kHz, Sidon + UTMOS-filtered)
Emolia-Thinking — Balanced Per-Dimension Bucket Subset A balanced, per-dimension bucket subset of VoiceNet/emolia-thinking, derived from that…