Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Dataset Card for Dataset Name We introduce the Egocentric Video Understanding Dataset (EVUD), an instruction-tuning dataset for training VLMs on video captioning and question answering tasks specific to egocentric videos. Dataset Details Dataset Description AI personal assistants deployed via robots or wearables require embodied understanding to collaborate with humans effectively. However, current Vision-Language Models (VLMs) primarily focus on third-person… See the full description on the dataset page:
Source: Hugging Face Hub (AlanaAI/EVUD). Metadata imported from the dataset’s Hub tags.