Skip to content
Advertisement

Video

706 results

Video

stream-data

Streaming Video Dataset Description A consolidated collection of video datasets for streaming video understanding research, including temporal…

100K–1M·CC-BY
MultimodalTextVideo

UrbanVideo-Bench

ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…

1K–10K·MIT·Parquet
MultimodalTextVideo

heico-focus-vqa

HeiCo-FOCUS (Beta release) A clinically grounded dataset for long-context video understanding in minimally invasive surgery. 📄 Paper  • …

10K–100K·CC-BY-NC-SA·Parquet
Video

LLaVA-Video-large-swift

Dataset Card LLaVA-Video-medium-swift A subset of LLaVA-Video-178K for educational purposes to learn how to fine-tune video models.

<1K·Apache-2.0
ImageMultimodalVideo

MindCraft

MindCraft Benchmark for spatio-temporal reasoning during vision-and-language navigation, used with LASAR. Layout mindcraft/ 000001/ data.npz 1.mp4…

1K–10K·Apache-2.0·Images (folder)
AudioMultimodalVideo

MRSDrama

ISDrama: Immersive Spatial Drama Generation through Multimodal Prompting Yu Zhang , Wenxiang Guo , Changhao Pan , Zhiyuan…

<1K·CC-BY-NC-SA
MultimodalTabularText

Talker-T2AV-Data

Talker-T2AV-Data Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling Paper (arXiv 2604.23586) · Code (GitHub) · Model ·…

100K–1M·Apache-2.0·CSV
Video

DAVIS-2017-480p-mp4

DAVIS 2017 — 480p MP4 Sequences Pre-encoded MP4 versions of all 90 sequences from the DAVIS 2017 trainval…

<1K·CC-BY-NC
Video

EchoNet-Dynamic-unzipped

EchoNet-Dynamic (Unzipped Version) This repository contains the unzipped video files of the EchoNet-Dynamic dataset. 🔗 Related Links Official…

10K–100K·MIT
ImageMultimodalVideo

IITKGP_Fence_dataset

IITKGP Fence dataset Overview The IITKGP Fence dataset is designed for tasks related to fence-like occlusion detection, defocus…

1K–10K·Apache-2.0·Images (folder)
ImageMultimodalVideo

TopAir

TopAir: Light-Weight Aerial dataset for Monocular Depth Estimation and Semantic Segmentation Curated by: Yara AlaaEldin Shared by: MaLga…

10K–100K·Images (folder)
Advertisement