BEAF
BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…
672 results
BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…
BoundingDocs 🔍 The largest spatially-annotated dataset for Document Question Answering Dataset Description BoundingDocs is a unified dataset for…
ICDAR2019's Scanned Receipts OCR and Information Extraction (SROIE)
PlantInquiryVQA — Thinking Like a Botanist
QCalEval Dataset Description The dataset contains scientific plots from quantum computing calibration experiments, paired with vision-language…
Molmo2 PointArena SFT Data (LAION + Molmo-7B-D)
HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence,…
Lost in Time: A New Temporal Benchmark for Video LLMs Daniel Cores , Michael Dorkenwald , Manuel Mucientes,…
DenseFusion-1M for comprehensive image descriptions