Galaxea-Open-World-Dataset
Galaxea Open-World Dataset Key Features 500+ hours of real-world mobile manipulation data. All data collected using one uniform…
293 results
Galaxea Open-World Dataset Key Features 500+ hours of real-world mobile manipulation data. All data collected using one uniform…
PROCESS-2: Speech Dataset for Early Cognitive Impairment Detection
Reazon Speech v2 DENOISED Same Speaker Pair prj-beatrice/reazon-speech-v2-denoised-titanet-embeddings のコサイン類似度の総和がなるべく大きくなるように…
Bo du lieu giong noi tieng Viet
Dataset Card for μ-Bench (Leaderboard Code) μ-Bench is a multilingual transcription benchmark built from real customer-service phone conversations.…
Moody Girl: Emotion-Tagged DailyTalk Dataset This dataset is an emotion-tagged version of the DailyTalk dataset, enhanced with emotion…
Peter Griffin's Emotion-Tagged DailyTalk Dataset Hehehehehe! Hey Lois, look! I made a dataset! This is an emotion-tagged version…
CV22-Sidon Overview This dataset hosts a release of Mozilla Common Voice 22 restored with the Sidon speech restoration…
Streaming ASR Dataset This dataset is designed for training real-time (streaming) ASR models, with a focus on handling…
IndicVoices: Towards building an Inclusive Multilingual Speech Dataset for Indian Languages Updates 23 December 2025 We now have…
Crowd Whatsapp Yiddish - Source Dataset
Knesset Plenum Source Dataset
Synthetic ASR data — hi Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…
Multilingual Speaker Diarization Dataset This dataset contains synthetic multilingual speaker diarization data with Hindi, English, and Punjabi audio…
Synthetic ASR data — zh Generated by Valsea-ASR/synthetic-data-pipeline. Audio is synthetic (TTS), targeted as training data for downstream…