Skip to content
Advertisement

Parquet

1,501 results

ImageMultimodalText

BEAF

BEAF: Before-After Changes for Hallucination Evaluation BEAF is a benchmark for evaluating object hallucination in vision-language models using…

10K–100K·CC-BY·Parquet
ImageMultimodalText

MultiChartQA

MultiChartQA This repository contains the questions and answers for our Multi-chart Benchmark. At present, only the data is…

1K–10K·CC-BY-NC·Parquet
ImageMultimodalText

DocVQA-2026

DocVQA 2026 ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains Building upon previous DocVQA benchmarks, this…

<1K·Parquet
ImageMultimodalText

GQA-ru

GQA-ru This is a translated version of original GQA dataset and stored in format supported for lmms-eval pipeline.…

10K–100K·Apache-2.0·Parquet
ImageMultimodalText

VLFeedback

Dataset Card for VLFeedback Homepage: Repository: Paper: Dataset Summary VLFeedback is a large-scale vision-language preference dataset, annotated by…

10K–100K·Parquet
ImageMultimodalText

BoundingDocs

BoundingDocs 🔍 The largest spatially-annotated dataset for Document Question Answering Dataset Description BoundingDocs is a unified dataset for…

10K–100K·CC-BY·Parquet
MultimodalTextVideo

UrbanVideo-Bench

ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…

1K–10K·MIT·Parquet
Advertisement