Visual Question Answering
339 results
LVOmniBench (QA repackaged for lmms-eval)
LVOmniBench (QA repackaged for lmms-eval)
FineBench
Dataset Card for FineBench FineBench is a large-scale, multiple-choice Video Question Answering (VQA) dataset designed specifically to evaluate…
VideoKR-Eval
VideoKR-Eval 📄 ArXiv | 💻 Code | 🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…
llava-en-zh-300k
This dataset is composed by 150k examples of English Visual Instruction Data from LLaVA. 150k examples of English…
FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal…
MVBench
MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…
pixmo-docs
PixMo-Docs We now recommend using CoSyn-400k and CoSyn-point over these datasets. They are improved versions with more images…
Visual Intelligence Leaderboard
Visual Intelligence Leaderboard
Nornikel Metallurgy VL Dataset
Nornikel Metallurgy VL Dataset
PhyX
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…
OpenDocVQA
Dataset Card for OpenDocVQA This is a training and evaluation QA data file for VDocRAG, a new RAG…
VLAA-Thinking
SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models 🌐 Project Page • 📄…