English
1,233 results
M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
SDG-30K
SDG-30K — Structured Defect Grounding Dataset A 30,000-image dataset for structured defect grounding in text-to-image generations. Each image…
GenEvolve-Data-Bench
GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose…
PlantInquiryVQA — Thinking Like a Botanist
PlantInquiryVQA — Thinking Like a Botanist
QCalEval
QCalEval Dataset Description The dataset contains scientific plots from quantum computing calibration experiments, paired with vision-language…
LVOmniBench (QA repackaged for lmms-eval)
LVOmniBench (QA repackaged for lmms-eval)
FineBench
Dataset Card for FineBench FineBench is a large-scale, multiple-choice Video Question Answering (VQA) dataset designed specifically to evaluate…
VideoKR-Eval
VideoKR-Eval 📄 ArXiv | 💻 Code | 🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…
FineReason-1.8M-Qwen3-VL-235B-Thinking
MMFineReason Closing the Multimodal Reasoning Gap via Open Data-Centric Methods Average score across mathematical reasoning and multimodal…
MVBench
MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…
Visual Intelligence Leaderboard
Visual Intelligence Leaderboard
PhyX
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…
OpenDocVQA
Dataset Card for OpenDocVQA This is a training and evaluation QA data file for VDocRAG, a new RAG…