Skip to content
Advertisement

Visual Question Answering

339 results

ImageMultimodalText

ChartVerse-SFT-1.8M

ChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…

1M–10M·Apache-2.0·Parquet
ImageMultimodalText

MMBench-ru

MMBench-ru This is a translated version of original MMBench dataset and stored in format supported for lmms-eval pipeline.…

1K–10K·Apache-2.0·Parquet
Text

pixelrag-tiles

PixelRAG tile corpus Rendered screenshot tiles for PixelRAG, a visual retrieval-augmented-generation system that retrieves over page images instead…

>1B·CC-BY-SA
ImageMultimodalTabular

SLAKE

Dataset Info: SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering ISBI 2021 oral Project Page: click…

10K–100K·CC-BY·JSON
ImageMultimodalText

ship-dataset

ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on…

10K–100K·CC-BY·JSON
ImageMultimodalText

Obshazard-bench

ObsCrisis-Bench A multimodal benchmark for evaluating large vision-language models on extreme weather event analysis tasks. Dataset Description…

<1K·MIT
Video

stream-data

Streaming Video Dataset Description A consolidated collection of video datasets for streaming video understanding research, including temporal…

100K–1M·CC-BY
ImageMultimodalText

CoSyn-400K

CoSyn-400k CoSyn-400k is a collection of synthetic question-answer pairs about very diverse range of computer-generated images. The data…

100K–1M·ODC-BY·Parquet
MultimodalTextVideo

UrbanVideo-Bench

ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…

1K–10K·MIT·Parquet
Advertisement