VideoKR-Eval
VideoKR-Eval 📄 ArXiv | 💻 Code | 🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…
634 results
VideoKR-Eval 📄 ArXiv | 💻 Code | 🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…
Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…
Video-opd-Dataset A video temporal grounding dataset with 2,500 samples sourced from TimeLens-100K. Dataset Description This dataset contains video…
Dataset Card for VideoMarathon VideoMarathon is a large-scale long video instruction-following dataset with a total duration of approximately…
Optimism Bias World Model Benchmark — Training Data Training data for VLM-as-judge models evaluating world model predictions. Structure…
MVTamperBench Dataset Overview MVTamperBenchEnd is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…
ICML 2026 Yiming Zhang1 , Jiacheng Chen1 , Jiaqi Tan1, Yongsen Mao2, Wenhu Chen3, Angel X. Chang1,4 1…
AVQA (Audio-Visual QA) — Videos + Annotations
Dataset Card for LongVALE Uses This dataset is designed for training and evaluating models on omni-modal (vision-audio-language-event) fine-grained…
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos Zongxia Li , Xiyang Wu , Guangyao Shi, Yubin…
Unify-OmniBench 统一格式的多模态评测数据集,由 Unify-OmniBench 框架转换生成。 包含四个 benchmark,在 Dataset Viewer 右上角下拉框切换。 数据概览 Config (bench) 题目数 模态 媒体 daily omni ~1197…
Lost in Time: A New Temporal Benchmark for Video LLMs Daniel Cores , Michael Dorkenwald , Manuel Mucientes,…
LapChole-FOCUS-VQA A clinically grounded benchmark for long-context video understanding in minimally invasive surgery. 💻 Code • 🏆…
V1: Toward Multimodal Reasoning by Designing Auxiliary Tasks 🚀 Toward Multimodal Reasoning via Unsupervised Task -- Future Prediction…
Multimodal Video QA Dataset This dataset contains challenging video question-answering tasks that require understanding both visual and audio…
MVTamperBench Dataset Overview MVTamperBenchStart is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…