Skip to content
Advertisement

Datasets

3,216 results

AudioMultimodalText

IndicTTS_Tamil

Tamil Indic TTS Dataset This dataset is derived from the Indic TTS Database project, specifically using the Tamil…

1K–10K·CC-BY·Parquet
AudioMultimodalText

Ace-Taffy-voice

关注永雏塔菲喵,关注永雏塔菲谢谢喵 永雏塔菲语音数据集 数据来自永雏塔菲直播,使用 silero vad + whisper-large-v3-trubo 粗略标注后,人工修正。 微调示例代码:

1K–10K·MIT·Parquet
AudioMultimodalText

dowis

Do What I Say (DOWIS): A Spoken Prompt Dataset for Instruction-Following NEW DOWIS now also contains spoken and…

1K–10K·CC-BY·Parquet
AudioMultimodalText

UltraVoice

UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack…

100K–1M·MIT·JSON
AudioMultimodalText

MiscSpeech-ja

MiscSpeech-ja This dataset comprises audio and corresponding transcripts collected from a diverse range of YouTube videos and Podcasts.

10K–100K·CC-BY-SA·Parquet
Audio

turkish-tts-combined-raw

Türkçe TTS Birleşik Veri Seti 7 farklı açık kaynak Türkçe TTS veri setinin birleşimi. ~81,500 örnek 24kHz SNAC…

10K–100K·CC-BY-SA
MultimodalTabularText

mls-annotated

Dataset Card for Annotations of non English MLS This dataset consists in annotations of a the Non English…

1M–10M·CC-BY·Parquet
AudioMultimodalText

ne-tts-nnp

NE-TTS Wancho (nnp) Cleaned TTS dataset for Wancho (nnp), a North East Indian language. Derived from the Vaani…

1K–10K·CC-BY·Parquet
Advertisement