ESpeech datasets annotate by Balalaika
ESpeech datasets annotate by Balalaika
2,368 results
ESpeech datasets annotate by Balalaika
Lahgtna Levantine TTS — Synthetic Levantine Arabic & Code-Switching
Common Voice 13 French (Phonemized & Curated)
Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…
Kahwa Postcast Arabic Speech Dataset
VOX-DUB is a human-based benchmark for evaluating AI dubbing systems.It includes: Audio fragments with original speech from real…
wuwa-voice-EN wuwa voice EN is a dataset of voice line from Wuthering Waves Attribute Value Language English Total…
Syrian Postcast Arabic Speech Dataset
InstructTTSEval InstructTTSEval is a comprehensive benchmark designed to evaluate Text-to-Speech (TTS) systems' ability to follow complex…
Live ATC Europe — audio + transcriptions Enregistrements live d'air traffic control (ATC) européen, capturés depuis LiveATC.net, segmentés…
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang , Wenxiang Guo , Changhao Pan , Zhiyuan…
Sagalee: An Open Source ASR Dataset for Oromo Language
Ghana TTS Navigation Corpus (Dagbani)
SynDataLab/Irodori-Ja-Spk1-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…
Contribution This dataset is a processed version of the original LibriTTS-P dataset, optimized for use on the Hugging…