1K–10K
598 results
VideoHallu
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos Zongxia Li , Xiyang Wu , Guangyao Shi, Yubin…
Unify-OmniBench
Unify-OmniBench 统一格式的多模态评测数据集,由 Unify-OmniBench 框架转换生成。 包含四个 benchmark,在 Dataset Viewer 右上角下拉框切换。 数据概览 Config (bench) 题目数 模态 媒体 daily omni ~1197…
MathVerse
Dataset Card for MathVerse Dataset Description Paper Information Dataset Examples Leaderboard Citation Dataset Description The capabilities of…
HoloCount
HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence,…
TVBench
Lost in Time: A New Temporal Benchmark for Video LLMs Daniel Cores , Michael Dorkenwald , Manuel Mucientes,…
MathVerse-lmmseval
Dataset Card for MathVerse This is the version for lmms-eval. This shares the same data with the official…
full-modality-bench
Multimodal Video QA Dataset This dataset contains challenging video question-answering tasks that require understanding both visual and audio…
ChartQAPro
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering 🤗Dataset 🖥️Code 📄Paper The abstract of the…
MathCanvas-Bench
MathCanvas-Bench 🚀 Data Usage from datasets import load dataset dataset = load dataset("shiwk24/MathCanvas-Bench") print(dataset) 📖 Introduction…
VideoThinkBench
CVPR 2026 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 🎊 News 2026.02 🔥🔥Our work…