VideoChat3-LV116k
VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video…
1,940 results
VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "fps": 30, "features": { "observation.state":…
Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News…
A dataset for "EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery"…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "Franka", "total episodes":…
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Project Page Paper GitHub Repository Deform360 is a…
NeurIPS D&B 2024 Spotlight ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation If you like our…
Audio2Tool — Spoken Tool-Calling Benchmark
Suno AI Music Dataset — Multi-Genre Curated
Dataset Card for LibriTTS LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech…
PROCESS-2: Speech Dataset for Early Cognitive Impairment Detection
FalAR FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the…
The dataset is available under the terms of the Creative Commons Attribution Non-Commercial license. K. J. Piczak. ESC:…
Russian voices for train AI. ♀ Male and ♂ Female Male voices - 497 pcs. Female voices -…
EuroSpeech 24 kHz Dataset Dataset Description EuroSpeech is a large-scale multilingual speech corpus containing high-quality aligned parliamentary…
Multilingual MFA-Aligned Speech Dataset