UrbanVideo-Bench
ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…
2,147 results
ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…
ICDAR2019's Scanned Receipts OCR and Information Extraction (SROIE)
M3CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
SDG-30K — Structured Defect Grounding Dataset A 30,000-image dataset for structured defect grounding in text-to-image generations. Each image…
GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose…
DCVLM-Baseline (200B tokens)
PlantInquiryVQA — Thinking Like a Botanist
QCalEval Dataset Description The dataset contains scientific plots from quantum computing calibration experiments, paired with vision-language…
LVOmniBench (QA repackaged for lmms-eval)
Dataset Card for FineBench FineBench is a large-scale, multiple-choice Video Question Answering (VQA) dataset designed specifically to evaluate…
VideoKR-Eval 📄 ArXiv | 💻 Code | 🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…