Skip to content
Advertisement

Hal-Eval: Hallucination Evaluation Benchmark A comprehensive benchmark for evaluating hallucination in vision-language models through caption comparison, from the paper “Hal-Eval: A Universal and Multi-Dimensional Benchmark for Hallucination Evaluation in Large Vision-Language Models.” Statistics Split Samples Images Source in domain 20,000 5,000 COCO val2014 out of domain 20,000 4,995 CC-SBU Total 40,000 9,995 Note: Out-of-domain samples reference… See the full description on the dataset page:

Source: Hugging Face Hub (MM-Hallu/Hal-Eval). Metadata imported from the dataset’s Hub tags.

Advertisement