Skip to content
Advertisement

CC-BY-NC-SA

149 results

AudioMultimodalText

Irodori-Ja-Spk2-10k

SynDataLab/Irodori-Ja-Spk2-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…

10K–100K·CC-BY-NC-SA·Parquet
AudioMultimodalVideo

MRSDrama

ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang , Wenxiang Guo , Changhao Pan , Zhiyuan…

<1K·CC-BY-NC-SA
AudioMultimodalText

Irodori-Ja-Spk1-10k

SynDataLab/Irodori-Ja-Spk1-10k 10,000 single-speaker conversational Japanese utterances synthesized with Irodori-TTS-500M-v2. Part of a 4-speaker…

10K–100K·CC-BY-NC-SA·Parquet
AudioMultimodalText

SynParaSpeech

Description Here is the SynParaSpeech dataset. SynParaSpeech is the first automated synthesis framework for constructing large-scale paralinguistic…

10K–100K·CC-BY-NC-SA·Parquet
AudioMultimodalText

Emilia-NV

NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at…

100K–1M·CC-BY-NC-SA·WebDataset
Multimodal

xview2-xbd

xView2 / xBD (mirror) Mirror of the xView2 / xBD building damage assessment dataset (parquet), used to train…

CC-BY-NC-SA
ImageMultimodalText

hescape-pyarrow

HESCAPE • PyArrow Format HESCAPE (H&E + Spatial Contrastive Pretraining Benchmark) is a large-scale benchmark for multimodal learning…

100K–1M·CC-BY-NC-SA·Parquet
Text

kits23

KiTS23 Dataset Dataset Description The KiTS23 dataset for kidney tumor segmentation. This dataset contains CT scans with dense…

<1K·CC-BY-NC-SA·JSON
Advertisement