MathCanvas-Bench
MathCanvas-Bench 🚀 Data Usage from datasets import load dataset dataset = load dataset("shiwk24/MathCanvas-Bench") print(dataset) 📖 Introduction…
1,233 results
MathCanvas-Bench 🚀 Data Usage from datasets import load dataset dataset = load dataset("shiwk24/MathCanvas-Bench") print(dataset) 📖 Introduction…
STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning…
CVPR 2026 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 🎊 News 2026.02 🔥🔥Our work…
CG-Bench Project Website: Repository: (includes running code) Summary We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question…
ChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…
PixelRAG tile corpus Rendered screenshot tiles for PixelRAG, a visual retrieval-augmented-generation system that retrieves over page images instead…
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on…
ObsCrisis-Bench A multimodal benchmark for evaluating large vision-language models on extreme weather event analysis tasks. Dataset Description…
Streaming Video Dataset Description A consolidated collection of video datasets for streaming video understanding research, including temporal…
SciReC: Diagnostic Evaluation of Multimodal Multi-Turn Relational Reasoning with Adaptive Interaction
ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…
Cambrian-Alignment Dataset Please see paper & website for more information: Overview Cambrian-Alignment is an question-answering alignment dataset…