Irodori-Ja-Spk2-10k
SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…
149 results
SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…
ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang , Wenxiang Guo , Changhao Pan , Zhiyuan…
SynDataLab/Irodori-Ja-Spk1-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…
Description Here is the SynParaSpeech dataset. SynParaSpeech is the first automated synthesis framework for constructing large-scale paralinguistic…
Korean Single Speaker Speech Dataset
Mimba PLT TTS Dataset (Plateau Malagasy Synthetic Speech)
NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at…
viVoice: Enabling Vietnamese Multi-Speaker Speech Synthesis
xView2 / xBD (mirror) Mirror of the xView2 / xBD building damage assessment dataset (parquet), used to train…
HESCAPE • PyArrow Format HESCAPE (H&E + Spatial Contrastive Pretraining Benchmark) is a large-scale benchmark for multimodal learning…
VALDO — Vascular Lesions Detection and Segmentation (MICCAI 2021)
Rock Segmentation Dataset for Excavator Rock-Picking Task
SegFly (RGB-T pairs FiftyOne subset)