HoloCount
HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence,…
1,940 results
HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence,…
Lost in Time: A New Temporal Benchmark for Video LLMs Daniel Cores , Michael Dorkenwald , Manuel Mucientes,…
LapChole-FOCUS-VQA A clinically grounded benchmark for long-context video understanding in minimally invasive surgery. 💻 Code • 🏆…
V1: Toward Multimodal Reasoning by Designing Auxiliary Tasks 🚀 Toward Multimodal Reasoning via Unsupervised Task -- Future Prediction…
StepCountQA-SFT StepCountQA-SFT is a multimodal supervised fine-tuning (SFT) dataset for visual object counting with step-by-step reasoning. Built…
Dataset Card for MathVerse This is the version for lmms-eval. This shares the same data with the official…
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning Repo: Paper: Introduction We introduce MathCoder-VL, a series…
ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering 🤗Dataset 🖥️Code 📄Paper The abstract of the…
MVTamperBench Dataset Overview MVTamperBenchStart is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…
DenseFusion-1M for comprehensive image descriptions
Viva Mais Synthetic PT-BR Travel Document Vision Dataset
MathCanvas-Bench 🚀 Data Usage from datasets import load dataset dataset = load dataset("shiwk24/MathCanvas-Bench") print(dataset) 📖 Introduction…
STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning…
🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 44.7K…