Skip to content
Advertisement

VisualBench: Temporal Video Understanding Benchmark 1600 synthetic video QA pairs designed to be 100% non-text-answerable (NTA). Every question requires watching the video’s temporal evolution — no single frame, no text-only shortcut can reveal the answer. Overview Property Value Total QAs 1600 Categories 16 Videos per category 100 Answer distribution 20% A, 20% B, 20% C, 20% D, 20% E Video format MP4, 720p30 Generation Manim Community v0.20.0… See the full description on the dataset page:

Source: Hugging Face Hub (AgPerry/VisualBench). Metadata imported from the dataset’s Hub tags.

Advertisement