Datasets
464 results
Uzbek YouTube Speech Dataset
Uzbek YouTube Speech Dataset
NaturalVoices Restored (48 kHz, Sidon + UTMOS-filtered)
NaturalVoices Restored (48 kHz, Sidon + UTMOS-filtered)
TTS Punctuation Vocal Commands FR
TTS Punctuation Vocal Commands FR
Google LATAM Spanish Female Speech with Normalized Boundary Silence
Google LATAM Spanish Female Speech with Normalized Boundary Silence
Open Bible Resources — African Languages
Open Bible Resources — African Languages
C3T
C3T: Cross-modal Capabilities Conservation Test Dataset Description C3T (Cross-modal Capabilities Conservation Test) is a benchmark for assessing the…
Dialogs: Expressive Conversational Russian Speech Corpus
Dialogs: Expressive Conversational Russian Speech Corpus
QuranLab — Qur’an Recitation Audio: Reference Manifest & CC-BY Word Timing
QuranLab — Qur'an Recitation Audio: Reference Manifest & CC-BY Word Timing
everyayah-phonemes
Phoneme-labelled Quran Datatset This dataset contains recitations from 45 professional Quran reciters, sourced from EveryAyah and QUL. The…
emolia-thinking-balanced-buckets
Emolia-Thinking — Balanced Per-Dimension Bucket Subset A balanced, per-dimension bucket subset of VoiceNet/emolia-thinking, derived from that…
viVoice: Enabling Vietnamese Multi-Speaker Speech Synthesis
viVoice: Enabling Vietnamese Multi-Speaker Speech Synthesis
BibleMMS
The Dataset associated with the Paper "Meta Learning Text-to-Speech Synthesis in over 7000 Languages" by Florian Lux, Sarina…