STRIDE-QA-Dataset
STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning…
2,368 results
STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning…
🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 44.7K…
CVPR 2026 Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm 🎊 News 2026.02 🔥🔥Our work…
CG-Bench Project Website: Repository: (includes running code) Summary We introduce CG-Bench, a groundbreaking benchmark for clue-grounded question…
ChartVerse-SFT-1800K is an extended large-scale chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…
MMBench-ru This is a translated version of original MMBench dataset and stored in format supported for lmms-eval pipeline.…
Dataset Info: SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering ISBI 2021 oral Project Page: click…
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on…
ObsCrisis-Bench A multimodal benchmark for evaluating large vision-language models on extreme weather event analysis tasks. Dataset Description…
SciReC: Diagnostic Evaluation of Multimodal Multi-Turn Relational Reasoning with Adaptive Interaction
CoSyn-400k CoSyn-400k is a collection of synthetic question-answer pairs about very diverse range of computer-generated images. The data…