whisper-dataset-ytb-uk
Dataset Card for Dataset Name This dataset is collected from youtube.
2,147 results
Dataset Card for Dataset Name This dataset is collected from youtube.
The "Thorsten-Voice" dataset This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of…
Cantonese Audio TTS Dataset This dataset contains alvanlii/cantonese-radio, alvanlii/cantonese-youtube, plus a dataset of equal size. It is catered…
Mostafa Mahmoud Arabic Speech Dataset
Tamil Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Tamil…
关注永雏塔菲喵,关注永雏塔菲谢谢喵 永雏塔菲语音数据集 数据来自永雏塔菲直播,使用 silero vad + whisper-large-v3-trubo 粗略标注后,人工修正。 微调示例代码:
Deep Confessions Podcast Arabic Speech Dataset
Do What I Say (DOWIS): A Spoken Prompt Dataset for Instruction-Following NEW DOWIS now also contains spoken and…
UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack…
MiscSpeech-ja This dataset comprises audio and corresponding transcripts collected from a diverse range of YouTube videos and Podcasts.
Mimba PLT TTS Dataset (Plateau Malagasy Synthetic Speech)
Dataset Card for Annotations of non English MLS This dataset consists in annotations of a the Non English…
Risale-i Nur Sesli Külliyat — Risale-i Nur Audio Corpus