Parquet
1,501 results
French_game_voice
French Game Voice Dataset Dataset of 100k+ cleaned audio samples of French video game voices with transcriptions. Features…
irodori-refs-10k-v2
Irodori TTS Reference Voices v2 (10K) 10,000 reference voices generated with Aratako/Irodori-TTS-500M-v2-VoiceDesign (no ref=True) using a richer…
bengali-tts-missing-v1
Bengali TTS — Missing Rows This dataset contains the rows from rwd51/bengali-tts-combined that are not present in the…
IndicTTS_Bengali
Bengali Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Bengali…
ivrit.ai – Knesset Plenums Whisper Training
ivrit.ai - Knesset Plenums Whisper Training
Eka Medical Asr Sample Noise Eval Dataset
Eka Medical Asr Sample Noise Eval Dataset
emova-sft-speech-231k
EMOVA-SFT-Speech-231K 🤗 EMOVA-Models 🤗 EMOVA-Datasets 🤗 EMOVA-Demo 📄 Paper 🌐 Project-Page 💻 Github 💻 EMOVA-Speech-Tokenizer-Github Overview…
Habibi
A systematic and standardized benchmark for the multi-dialect Arabic zero-shot TTS task Paper: "Habibi: Laying the Open-Source Foundation…
InfoRe Technology public dataset №1
InfoRe Technology public dataset №1
Emolia · Filtered · NanoCodec (FSQ) Tokens
Emolia · Filtered · NanoCodec (FSQ) Tokens
ESpeech datasets annotate by Balalaika
ESpeech datasets annotate by Balalaika
Lahgtna Levantine TTS — Synthetic Levantine Arabic & Code-Switching
Lahgtna Levantine TTS — Synthetic Levantine Arabic & Code-Switching
Common Voice 13 French (Phonemized & Curated)
Common Voice 13 French (Phonemized & Curated)
bashkort_tts_dataset
Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…
Kahwa Postcast Arabic Speech Dataset
Kahwa Postcast Arabic Speech Dataset