Skip to content
Advertisement

English

1,233 results

MultimodalTextVideo

Vript

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 12K…

100K–1M·JSON
ImageMultimodalVideo

GOKU-2M

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing GOKU-2M is a large-scale, unified instruction-based…

1M–10M·CC-BY-NC
MultimodalTextVideo

finevideo

FineVideo FineVideo Description Dataset Explorer Revisions Dataset Distribution How to download and use FineVideo Using datasets Using huggingface…

10K–100K·Other·Parquet
MultimodalTabularText

full-modality-data

Full Modality Dataset Statistics Video Statistics Total Videos: 28,472 Total Duration: 1422.33 hours Average Duration: 179.84 seconds Median…

1M–10M·MIT·Parquet
Video

RoVid-X

Rethinking Video Generation Model for the Embodied World If you like our project, please give us a star…

>1B·CC-BY
Video

MVLU

MLVU: Multi-task Long Video Understanding Benchmark This repo contains the annotation data and evaluation code for the paper…

CC-BY-NC-SA
MultimodalTextVideo

VideoChat3-LV116k

VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video…

1K–10K·Apache-2.0·JSON
Video

FakeParts_Legacy

FakeParts: A New Family of AI-Generated DeepFakes Abstract We introduce FakeParts, a new class of deepfakes characterized by…

<1K·CC0
Advertisement