X-WAM-RoboTwin
X-WAM Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Dataset Summary This is the RoboTwin…
595 results
X-WAM Unified 4D World Action Modeling from Video Priors with Asynchronous Denoising Dataset Summary This is the RoboTwin…
Dataset Card for LibriTTS LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech…
FalAR FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the…
Tarteel AI - EveryAyah Dataset
Dataset Card for LibriTTS-R LibriTTS-R 1 is a sound quality improved version of the LibriTTS corpus ( which…
MRSAudio: A Large-Scale Multimodal Recorded Spatial Audio Dataset with Refined Annotations Humans rely on multisensory integration to perceive…
This repository contains the vocal bursts like giggling, laughter, shouting, crying, etc. from the following repository. We captioned…
shkolkovo-bobr.video-webinars-audio Dataset of audio of ≈2573 webinars from bobr.video with text transcription made with whisper and VAD. Webinars…
FreeSound.org LAION-640k Dataset (16 KHz)
قاعدة بيانات المعلم القرآنية هذه ال dataset هي جزء من مشروع الملم الرقرآني: quran-muaalem وهي تهدف لكشف أخاطاء…
JamendoMaxCaps Dataset JamendoMaxCaps is a large-scale dataset of over 362,000 instrumental tracks sourced from the Jamendo platform. It…
Streaming ASR Dataset This dataset is designed for training real-time (streaming) ASR models, with a focus on handling…