Skip to content
Advertisement

Datasets

3,216 results

AudioMultimodalText

bashkort_tts_dataset

Bashkort TTS Dataset The largest open dataset for speech synthesis in the Bashkir language — featuring multi-speaker recordings…

10K–100K·CC-BY·Parquet
AudioMultimodalText

VOX-DUB

VOX-DUB is a human-based benchmark for evaluating AI dubbing systems.It includes: Audio fragments with original speech from real…

1K–10K·Parquet
AudioMultimodalText

wuwa-voice-EN

wuwa-voice-EN wuwa voice EN is a dataset of voice line from Wuthering Waves Attribute Value Language English Total…

10K–100K·Parquet
AudioMultimodalText

InstructTTSEval

InstructTTSEval InstructTTSEval is a comprehensive benchmark designed to evaluate Text-to-Speech (TTS) systems' ability to follow complex…

1K–10K·MIT·Parquet
AudioMultimodalText

live-atc-europe

Live ATC Europe — audio + transcriptions Enregistrements live d'air traffic control (ATC) européen, capturés depuis LiveATC.net, segmentés…

100K–1M·Custom / Research-only·Parquet
AudioMultimodalVideo

MRSDrama

ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang , Wenxiang Guo , Changhao Pan , Zhiyuan…

<1K·CC-BY-NC-SA
AudioMultimodalText

Irodori-Ja-Spk1-10k

SynDataLab/Irodori-Ja-Spk1-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…

10K–100K·CC-BY-NC-SA·Parquet
Advertisement