Datasets
911 results
MusicCaps
Dataset Card for MusicCaps Dataset Summary The MusicCaps dataset contains 5,521 music examples, each of which is labeled…
enhanced-audiosnippets-long-2-8M
Enhanced Audiosnippets Long 2.8M Enhanced version of mitermix/audiosnippets long 2 8M with speech enhancement, emotion annotations, speaker…
Arabic Diacritized-Stem Lexicon
Arabic Diacritized-Stem Lexicon
Urdu-ONYX-WAV-kanade-Annotated
Urdu-ONYX-WAV-real-Annotated Enhanced version of Urdu-ONYX-WAV-real with phoneme annotations and Kanade tokenizer features. Dataset Statistics Total…
moss-character-voices-bestof64
MOSS Character Voices — Best-of-64 (Stage 2) Best-of-64 voice-acting takes from the 4.55B MOSS-TTS-Local voice-acting model…
Talker-T2AV-Data
Talker-T2AV-Data Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling Paper (arXiv 2604.23586) · Code (GitHub) · Model ·…
nutriderm-dataset
NutriDermAI Dataset Dataset for NutriDermAI — Multimodal AI System for Dermatology with ABCDE Explainability and VQA. M.Tech Thesis…
GeoMeld
🌍 GeoMeld Multi-Modal Earth Observation Dataset (WebDataset) GeoMeld is a large-scale multi-modal remote sensing dataset introduced in our…
Synthetic Inline Holographical Images v3 (224px Highly Diverse)
Synthetic Inline Holographical Images v3 (224px Highly Diverse)
Perle AI Multi-phase CECT and CT with Radiology Reports
Perle AI Multi-phase CECT and CT with Radiology Reports
ACDC (Cardiac Cine-MRI)
ACDC (Cardiac Cine-MRI)
ACDC (Cardiac Cine-MRI)
ACDC (Cardiac Cine-MRI)
CTSpinoPelvic1K
CTSpinoPelvic1K A fused spine + pelvis 3D CT segmentation dataset built by patient-level crosswalk between three public sources:…
TerraCoT
TerraCoT TerraCoT is a remote-sensing chain-of-thought VQA dataset in which every SEG token in an answer is grounded…
ZJU Eye-Pretrain (Private Shanghai Topcon + 56 public cohorts, 8 modalities)
ZJU Eye-Pretrain (Private Shanghai Topcon + 56 public cohorts, 8 modalities)