Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
EatVid-Bench: A Multimodal Fine-Grained Eating Behavior Video Dataset Dataset Description EatVid-Bench is a multimodal benchmark dataset for evaluating video understanding models on fine-grained eating behavior analysis. It contains anonymized long-form eating videos and question-answer annotations covering food recognition, utensil perception, action understanding, and temporal reasoning. The dataset is designed to support research on long-form video… See the full description on the dataset page:
Source: Hugging Face Hub (linlingw/EatVid-Bench). Metadata imported from the dataset’s Hub tags.