Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…
ACL’25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains the dataset introduced in the paper, consisting of two parts: 5k+ multiple-choice question-answering (MCQ) data and 1k+ video clips. Arxiv: Project: Code: Dataset Description The… See the full description on the dataset page:
Source: Hugging Face Hub (EmbodiedCity/UrbanVideo-Bench). Metadata imported from the dataset’s Hub tags.