Skip to content
Advertisement

Multimodal

2,368 results

MultimodalTextVideo

CSL-News

Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News…

100K–1M·CC-BY-NC·JSON
MultimodalTextVideo

EyePCR

A dataset for "EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery"…

10K–100K·CC-BY·JSON
MultimodalTextVideo

deform360

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Project Page Paper GitHub Repository Deform360 is a…

MIT
MultimodalTextVideo

ChronoMagic-Bench

NeurIPS D&B 2024 Spotlight ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation If you like our…

1K–10K·CC-BY·CSV
AudioMultimodalText

libritts

Dataset Card for LibriTTS LibriTTS is a multi-speaker English corpus of approximately 585 hours of read English speech…

100K–1M·CC-BY·Parquet
AudioMultimodalText

FalAR

FalAR FalAR is a large-scale, speaker-annotated European Portuguese speech corpus built from recordings of parliamentary sessions of the…

100K–1M·CC-BY·Parquet
AudioMultimodalText

esc50

The dataset is available under the terms of the Creative Commons Attribution Non-Commercial license. K. J. Piczak. ESC:…

1K–10K·Parquet
Advertisement