Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…
BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using before-after image manipulation pairs. 26,064 QA pairs over 2,223 images (500 original COCO images + 1,723 manipulated images) with POPE-style yes/no questions. Fields Field Description image The image (original COCO or manipulated) question POPE-style question: “Is there a/an {object} in the image?” gt Ground truth… See the full description on the dataset page:
Source: Hugging Face Hub (MM-Hallu/BEAF). Metadata imported from the dataset’s Hub tags.