Skip to content
Advertisement
Video

RoboFine-Bench

RoboFine-Bench

RoboFine-Bench A Fine-Grained Robotic Video Understanding Benchmark RoboFine-Bench is a benchmark for evaluating whether Vision-Language Models (VLMs) can capture execution-level details of robot manipulation — going beyond coarse task recognition to understand how a robot performs a task. It was introduced in the paper FineVLA: Fine-Grained Instruction Alignment for Steerable Vision-Language-Action Policies and is part of the FineVLA framework for fine-grained instruction… See the full description on the dataset page:

Source: Hugging Face Hub (xlangai/RoboFine-bench). Metadata imported from the dataset’s Hub tags.

Advertisement