Skip to content
Advertisement

Video

706 results

Video

hoigen-filtered-videos

HOIGen Filtered Videos Dataset This dataset contains 28562 filtered videos from the HOIGen-1M dataset based on the allowlist.…

1K–10K·MIT
MultimodalTextVideo

M3arsSynth

Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions M3arsSynth Dataset Summary M3arsSynth is a large-scale,…

CC-BY
3D / Point CloudMultimodalVideo

DSI-Bench

DSI-Bench: A Benchmark for Dynamic Spatial Intelligence Paper Project Page Code Abstract Reasoning about dynamic spatial relationships is…

1K–10K·CC-BY
ImageMultimodalText

RoboCerebra

Overview RoboCerebra is a large-scale benchmark designed to evaluate robotic manipulation in long-horizon tasks, shifting the focus of…

1K–10K·MIT·Parquet
MultimodalSensor / Time-seriesTabular

libero

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "panda", "total episodes":…

100K–1M·Apache-2.0·Parquet
MultimodalTextVideo

phyworldbench

PhyWorldBench This repository hosts the core assets of PhyWorldBench, the 1,050 JSON prompt files, the evaluation standards, and…

<1K·MIT·Parquet
MultimodalTextVideo

REVISOR-25k

REVISOR-25k A multi-task video understanding dataset for training video LLMs with reinforcement learning (GRPO). The dataset contains ~25k…

10K–100K·Apache-2.0·Parquet
Video

bridge_v2_lerobot_pathmask

PEEK VLM-Labeled BRIDGE v2 dataset This dataset contains the LeRobot-format BRIDGE-v2 dataset with paths and masks from the…

Apache-2.0
Video

OmniMMI

OmniMMI Paper: OmniMMI: A Comprehensive Multi-modal Interaction Benchmark in Streaming Video Contexts Code Dataset Description we introduce OmniMMI,…

1K–10K·MIT
Advertisement