Skip to content
Advertisement

Video

706 results

MultimodalTabularText

full-modality-data

Full Modality Dataset Statistics Video Statistics Total Videos: 28,472 Total Duration: 1422.33 hours Average Duration: 179.84 seconds Median…

1M–10M·MIT·Parquet
MultimodalSensor / Time-seriesTabular

pusht

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v2.0", "robot type": "unknown", "total episodes":…

10K–100K·MIT·Parquet
Video

RoVid-X

Rethinking Video Generation Model for the Embodied World If you like our project, please give us a star…

>1B·CC-BY
Video

MVLU

MLVU: Multi-task Long Video Understanding Benchmark This repo contains the annotation data and evaluation code for the paper…

CC-BY-NC-SA
MultimodalTextVideo

VideoChat3-LV116k

VideoChat3-LV116K VideoChat3-LV116K is the long-video instruction data used by VideoChat3. It is designed to complement short academic video…

1K–10K·Apache-2.0·JSON
Video

memo_data

MEMO Video Dataset This dataset contains a curated collection of human talking videos gathered from publicly accessible sources…

10K–100K·CC-BY
Video

FakeParts_Legacy

FakeParts: A New Family of AI-Generated DeepFakes Abstract We introduce FakeParts, a new class of deepfakes characterized by…

<1K·CC0
Video

siftformer2-data

SiftFormer2 Data K400 / SSv2 video classification 학습용 데이터. 구성 siftformer2-data/ ├── k400/ │ ├── train manifest.jsonl │…

10K–100K·MIT
MultimodalTextVideo

CSL-News

Summary This is the dataset proposed in our paper "Uni-Sign: Toward Unified Sign Language Understanding at Scale". CSL-News…

100K–1M·CC-BY-NC·JSON
Advertisement