BEAF
BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…
1,501 results
BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…
MultiChartQA This repository contains the questions and answers for our Multi-chart Benchmark. At present, only the data is…
DocVQA 2026 ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains Building upon previous DocVQA benchmarks, this…
GQA-ru This is a translated version of original GQA dataset and stored in format supported for lmms-eval pipeline.…
VisionArena-Bench: An automatic eval pipeline to estimate model preference rankings An automatic benchmark of 500 diverse user prompts…
Dataset Card for VLFeedback Homepage: Repository: Paper: Dataset Summary VLFeedback is a large-scale vision-language preference dataset, annotated by…
BoundingDocs 🔍 The largest spatially-annotated dataset for Document Question Answering Dataset Description BoundingDocs is a unified dataset for…
ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…