Skip to content
Advertisement

MIT

552 results

Audio

Urdu-Munch-Lina

Urdu-Munch-Lina Processed version of zuhri025/Urdu-Munch with LinaCodec encoding. Dataset Structure This dataset contains 2 batches of audio data…

10K–100K·MIT
AudioMultimodalText

InstructTTSEval

InstructTTSEval InstructTTSEval is a comprehensive benchmark designed to evaluate Text-to-Speech (TTS) systems' ability to follow complex…

1K–10K·MIT·Parquet
Audio

audio_data_russian_backup

Dataset Audio Russian Backup This is a backup dataset with Russian audio data, split into train 0 to…

100K–1M·MIT
AudioMultimodalText

Ace-Taffy-voice

关注永雏塔菲喵,关注永雏塔菲谢谢喵 永雏塔菲语音数据集 数据来自永雏塔菲直播,使用 silero vad + whisper-large-v3-trubo 粗略标注后,人工修正。 微调示例代码:

1K–10K·MIT·Parquet
AudioMultimodalText

UltraVoice

UltraVoice: Scaling Fine-Grained Style-Controlled Speech Conversations for Spoken Dialogue Models 📝 Abstract Spoken dialogue models currently lack…

100K–1M·MIT·JSON
AudioMultimodalText

BibleMMS

The Dataset associated with the Paper "Meta Learning Text-to-Speech Synthesis in over 7000 Languages" by Florian Lux, Sarina…

100K–1M·MIT·Parquet
ImageMultimodalText

IndustryShapes

IndustryShapes Project Page Paper IndustryShapes is a new benchmark dataset tailored for 6D object pose estimation in industrial…

10K–100K·MIT·Parquet
Advertisement