Skip to content
Advertisement

Text

2,175 results

AudioMultimodalText

dowis

Do What I Say (DOWIS): A Spoken Prompt Dataset for Instruction-Following NEW DOWIS now also contains spoken and…

1K–10K·CC-BY·Parquet
AudioMultimodalText

UltraVoice

UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack…

100K–1M·MIT·JSON
AudioMultimodalText

MiscSpeech-ja

MiscSpeech-ja This dataset comprises audio and corresponding transcripts collected from a diverse range of YouTube videos and Podcasts.

10K–100K·CC-BY-SA·Parquet
MultimodalTabularText

mls-annotated

Dataset Card for Annotations of non English MLS This dataset consists in annotations of a the Non English…

1M–10M·CC-BY·Parquet
AudioMultimodalText

ne-tts-nnp

NE-TTS Wancho (nnp) Cleaned TTS dataset for Wancho (nnp), a North East Indian language. Derived from the Vaani…

1K–10K·CC-BY·Parquet
AudioMultimodalText

vctk

VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero…

10K–100K·CC-BY·Parquet
AudioMultimodalText

OpenDialog

OpenDialog OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation…

100K–1M·CC-BY-NC·WebDataset
AudioMultimodalText

Emilia-NV

NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at…

100K–1M·CC-BY-NC-SA·WebDataset
AudioMultimodalText

spoken-magpie-ja

Spoken-magpie LLMの日本語Instruction Tuning用データllm-jp/magpie-sft-v1.0をCosyVoice2 TTSを使用して音声化した商用利用可能な日本語の音声言語モデルのSFT用データセットです。 ある程度の話者多様性を持つように生成されています。…

100K–1M·Apache-2.0·Parquet
Advertisement