Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence, requiring a…
HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence, requiring a seamless integration of fine-grained grounding and spatial reasoning. While Multimodal Large Language Models (MLLMs) have achieved remarkable success in qualitative scene understanding, their quantitative precision remains a significant bottleneck, often characterized by… See the full description on the dataset page:
Source: Hugging Face Hub (MM-MVR/HoloCount). Metadata imported from the dataset’s Hub tags.