Skip to content
Advertisement

Text-to-Audio

26 results

AudioMultimodalVideo

MRSDrama

ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang , Wenxiang Guo , Changhao Pan , Zhiyuan…

<1K·CC-BY-NC-SA
AudioMultimodalText

TV-44kHz-Full

The "Thorsten-Voice" dataset This truly open source (CC0 license) german (🇩🇪) voice dataset contains about 40 hours of…

10K–100K·CC0
AudioMultimodalText

vctk

VCTK This is a processed clone of the VCTK dataset with leading and trailing silence removed using Silero…

10K–100K·CC-BY·Parquet
ImageMultimodalText

AVGen-Bench

AVGen-Bench Generated Videos Data Card Overview This data card describes the generated audio-video outputs stored directly in the…

1K–10K·MIT·Parquet
Advertisement