Skip to content
Advertisement

Datasets

1,940 results

ImageMultimodalText

HoloCount

HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence,…

1K–10K·CC-BY·Images (folder)
MultimodalTabularText

TVBench

Lost in Time: A New Temporal Benchmark for Video LLMs Daniel Cores , Michael Dorkenwald , Manuel Mucientes,…

1K–10K·CC-BY·JSON
MultimodalTextVideo

lapchole-focus-vqa

LapChole-FOCUS-VQA A clinically grounded benchmark for long-context video understanding in minimally invasive surgery. 💻 Code  •  🏆…

10K–100K·Custom / Research-only·Parquet
MultimodalTextVideo

V1-33K

V1: Toward Multimodal Reasoning by Designing Auxiliary Tasks 🚀 Toward Multimodal Reasoning via Unsupervised Task -- Future Prediction…

10K–100K·Apache-2.0·JSON
ImageMultimodalText

StepCountQA-SFT

StepCountQA-SFT StepCountQA-SFT is a multimodal supervised fine-tuning (SFT) dataset for visual object counting with step-by-step reasoning. Built…

1M–10M·ODC-BY·Parquet
ImageMultimodalText

ImgCode-8.6M

MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning Repo: Paper: Introduction We introduce MathCoder-VL, a series…

1M–10M·Apache-2.0·Parquet
Text

ChartQAPro

ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering 🤗Dataset 🖥️Code 📄Paper The abstract of the…

1K–10K·MIT·Parquet
MultimodalTabularText

MVTamperBenchStart

MVTamperBench Dataset Overview MVTamperBenchStart is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…

10K–100K·MIT·JSON
ImageMultimodalText

MathCanvas-Bench

MathCanvas-Bench 🚀 Data Usage from datasets import load dataset dataset = load dataset("shiwk24/MathCanvas-Bench") print(dataset) 📖 Introduction…

1K–10K·Apache-2.0·Parquet
ImageMultimodalText

STRIDE-QA-Dataset

STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning…

100K–1M·WebDataset
MultimodalTextVideo

Vript_Chinese

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 44.7K…

100K–1M·JSON
Advertisement