Skip to content
Advertisement

Text

2,175 results

MultimodalTextVideo

phyground

PhyGround Project Page GitHub Paper PhyGround is a criteria-grounded benchmark for evaluating physical reasoning in video generation. The…

<1K·JSON
MultimodalTextVideo

Vript

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 12K…

100K–1M·JSON
MultimodalTextVideo

finevideo

FineVideo FineVideo Description Dataset Explorer Revisions Dataset Distribution How to download and use FineVideo Using datasets Using huggingface…

10K–100K·Other·Parquet
MultimodalTabularText

full-modality-data

Full Modality Dataset Statistics Video Statistics Total Videos: 28,472 Total Duration: 1422.33 hours Average Duration: 179.84 seconds Median…

1M–10M·MIT·Parquet
MultimodalTextVideo

VideoChat3-LV116k

VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video…

1K–10K·Apache-2.0·JSON
MultimodalTextVideo

CSL-News

Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News…

100K–1M·CC-BY-NC·JSON
MultimodalTextVideo

EyePCR

A dataset for "EyePCR: A Comprehensive Benchmark for Fine-Grained Perception, Knowledge Comprehension and Clinical Reasoning in Ophthalmic Surgery"…

10K–100K·CC-BY·JSON
MultimodalTextVideo

deform360

Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Project Page Paper GitHub Repository Deform360 is a…

MIT
MultimodalTextVideo

ChronoMagic-Bench

NeurIPS D&B 2024 Spotlight ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation If you like our…

1K–10K·CC-BY·CSV
Advertisement