Skip to content
Advertisement

Multimodal

2,368 results

AudioMultimodalText

parczech4speech-segmented

ParCzech4Speech (Sentence-Segmented Variant) Dataset Summary ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based…

100K–1M·CC-BY·WebDataset
AudioMultimodalText

INTP

INTP: Intelligibility Preference Speech Dataset We establish a synthetic Intelligibility Preference Speech Dataset (INTP), including about 250K…

100K–1M·CC-BY-NC·Parquet
AudioMultimodalText

Pretraining-V1

Indic TTS Unified v1 A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This…

10M–100M·CC-BY·Parquet
ImageMultimodalText

emova-alignment-7m

EMOVA-Alignment-7M 🤗 EMOVA-Models 🤗 EMOVA-Datasets 🤗 EMOVA-Demo 📄 Paper 🌐 Project-Page 💻 Github 💻 EMOVA-Speech-Tokenizer-Github Overview…

1M–10M·Apache-2.0·Parquet
ImageMultimodalText

emova-sft-4m

EMOVA-SFT-4M 🤗 EMOVA-Models 🤗 EMOVA-Datasets 🤗 EMOVA-Demo 📄 Paper 🌐 Project-Page 💻 Github 💻 EMOVA-Speech-Tokenizer-Github Overview EMOVA-SFT-4M is…

1M–10M·Apache-2.0·Parquet
MultimodalTabularText

Talker-T2AV-Data

Talker-T2AV-Data Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling Paper (arXiv 2604.23586) · Code (GitHub) · Model ·…

100K–1M·Apache-2.0·CSV
ImageMultimodalTabular

nutriderm-dataset

NutriDermAI Dataset Dataset for NutriDermAI — Multimodal AI System for Dermatology with ABCDE Explainability and VQA. M.Tech Thesis…

1M–10M·CC-BY·Parquet
ImageMultimodalText

IndustryShapes

IndustryShapes Project Page Paper IndustryShapes is a new benchmark dataset tailored for 6D object pose estimation in industrial…

10K–100K·MIT·Parquet
Advertisement