CC-BY
672 results
FalAR-TTS
FalAR-TTS FalAR-TTS is a subset of the FalAR dataset ( tailored for speech synthesis in European Portuguese. To…
InfoRe Technology public dataset №2
InfoRe Technology public dataset №2
Hinglish Concatenated Audio Dataset
Hinglish Concatenated Audio Dataset
VoxCPM Ghana — Precomputed AudioVAE Latents
VoxCPM Ghana — Precomputed AudioVAE Latents
YodasSpeakerPool
Use this dataset in conjuction with: YodasSpeakerPool YodasSpeakerPool is a curated, richly-annotated multi-speaker dataset featuring 7,600 unique…
Character Voices — DramaBox-annotated Echo-TTS
Character Voices — DramaBox-annotated Echo-TTS
VLSP 2020 – VinAI – ASR challenge dataset
VLSP 2020 - VinAI - ASR challenge dataset
ne-tts-ccp
NE-TTS Chakma (ccp) Cleaned TTS dataset for Chakma (ccp), a North East Indian language. Derived from the Vaani…
linto-dataset-audio-ar-tn
LinTO DataSet Audio for Arabic Tunisian A collection of Tunisian dialect audio and its annotations for STT task…
Ganjoor Persian Poetry Recitations (Full)
Ganjoor Persian Poetry Recitations (Full)
Vāgdhenu — Sanskrit Chant Corpus
Vāgdhenu — Sanskrit Chant Corpus
Emolia-Thinking (VoiceNet balanced subset)
Emolia-Thinking (VoiceNet balanced subset)
linto-dataset-audio-ar-tn-augmented
LinTO DataSet Audio for Arabic Tunisian Augmented A collection of Tunisian dialect audio and its annotations for STT…
a novel large-scale Vietnamese speech corpus (LSVSC)
a novel large-scale Vietnamese speech corpus (LSVSC)
SimbaBench_dataset
SibmaBench Data Release & Benchmarking To evaluate your model on SimbaBench across all supported tasks (ASR, TTS, and…