Skip to content
Advertisement

CC-BY-NC

210 results

ImageMultimodalText

Kvasir-VQA-x1

Kvasir-VQA-x1 A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Kvasir-VQA-x1 on GitHub Original Image…

100K–1M·CC-BY-NC·Parquet
ImageMultimodalText

ChartGalaxy

ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation 🤗 Dataset 🖥️ Code 📄 Paper 🔥 News 2026.02…

1K–10K·CC-BY-NC·Parquet
AudioMultimodalText

OpenDialog

OpenDialog OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation…

100K–1M·CC-BY-NC·WebDataset
MultimodalTabularText

UniST

UniST This dataset contains UniST codec-token training data exported from local metadata and codec results. We train UniSS…

10M–100M·CC-BY-NC
Audio

LibriQuote

This repository contains the LibriQuote dataset, a speech dataset of fictional character utterances for expressive zero-shot speech synthesis.…

10M–100M·CC-BY-NC
Text

CapSpeech

CapSpeech DataSet used for the paper: CapSpeech: Enabling Downstream Applications in Style-Captioned Text-to-Speech Please refer to CapSpeech repo…

10M–100M·CC-BY-NC·Parquet
AudioMultimodalText

SaSLaW

This repository contains the data of SaSLaW corpus. You can download it via the following command: huggingface-cli download…

1K–10K·CC-BY-NC·Audio (folder)
AudioMultimodalText

nonverbalspeech38k

🎉 🎉 🎉 NonVerbalSpeech-38K: A Scalable Pipeline for Enabling Non-Verbal Speech Generation and Understanding The official repository for…

10K–100K·CC-BY-NC·Parquet
AudioMultimodalText

SpeechJudge-Data

SpeechJudge-Data: A Large-Scale Human Feedback Corpus for Speech Generation Introduction SpeechJudge-Data is a large-scale human feedback corpus of…

10K–100K·CC-BY-NC·Parquet
AudioMultimodalText

INTP

INTP: Intelligibility Preference Speech Dataset We establish a synthetic Intelligibility Preference Speech Dataset (INTP), including about 250K…

100K–1M·CC-BY-NC·Parquet
Advertisement