Skip to content
Advertisement

CC-BY

672 results

ImageMultimodalText

BEAF

BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…

10K–100K·CC-BY·Parquet
Image

CARV

CARV

<1K·CC-BY·Images (folder)
ImageMultimodalText

BoundingDocs

BoundingDocs 🔍 The largest spatially-annotated dataset for Document Question Answering Dataset Description BoundingDocs is a unified dataset for…

10K–100K·CC-BY·Parquet
ImageMultimodalText

QCalEval

QCalEval Dataset Description The dataset contains scientific plots from quantum computing calibration experiments, paired with vision-language…

<1K·CC-BY·Parquet
Image

PanoEnv

CVPR 2026 Highlight - PanoEnv-QA: A Large-Scale Geometry-Grounded Panoramic VQA Benchmark for 3D Spatial Intelligence 📖 Overview PanoEnv-QA…

10K–100K·CC-BY·Images (folder)
Text

MapTrace

MapTrace: A 2M-Sample Synthetic Dataset for Path Tracing on Maps Welcome to the MapTrace dataset! If you use…

10K–100K·CC-BY·Parquet
ImageMultimodalText

HoloCount

HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence,…

1K–10K·CC-BY·Images (folder)
MultimodalTabularText

TVBench

Lost in Time: A New Temporal Benchmark for Video LLMs Daniel Cores , Michael Dorkenwald , Manuel Mucientes,…

1K–10K·CC-BY·JSON
Advertisement