Skip to content
Advertisement

Visual Question Answering

339 results

MultimodalTabularText

MVTamperBench

MVTamperBench Dataset Overview MVTamperBench is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…

10K–100K·MIT
ImageMultimodalTabular

MVBench

MVBench Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded…

1K–10K·MIT·JSON
ImageMultimodalTabular

zendo-synthetic-data

Zendo Synthetic Visual Reasoning Dataset Synthetic Zendo-style scenes with associated rules and per-scene tensor representations. Each scene either…

10K–100K·CC-BY·Parquet
MultimodalTextVideo

VSI-Bench

Dataset arXiv Website Code VSI-Bench VSI-Bench-Debiased !IMPORTANT Nov. 7, 2025 UPDATE: This Dataset has been updated to include…

10K–100K·Apache-2.0·Parquet
ImageMultimodalTabular

MentalBlackboard

MentalBlackboard Benchmark This repository contains the MentalBlackboard benchmark with multiple tasks: - Prediction - Planning Dataset Sources

1K–10K·CC-BY·Parquet
MultimodalTextVideo

Vript

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 12K…

100K–1M·JSON
MultimodalTextVideo

finevideo

FineVideo FineVideo Description Dataset Explorer Revisions Dataset Distribution How to download and use FineVideo Using datasets Using huggingface…

10K–100K·Other·Parquet
Video

ParaVT-Source

ParaVT-Source Source media archives for the ParaVT training corpus. Pair this repository with the annotations in ParaVT/ParaVT-Parquet. Overview…

100K–1M·Apache-2.0
Advertisement