Skip to content
Advertisement

English

1,233 results

Video

ViCA-322K

ViCA-322K: A Dataset for Visuospatial Cognition in Real-World Indoor Videos Quickstart You can load our dataset using the…

Image

VisuLogic

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models A Challenging Visual-centric Benchmark for Evaluating…

1K–10K·Apache-2.0·Images (folder)
ImageMultimodalText

SDG-30K

SDG-30K — Structured Defect Grounding Dataset A 30,000-image dataset for structured defect grounding in text-to-image generations. Each image…

10K–100K·CC-BY-NC·JSON
ImageMultimodalText

GenEvolve-Data-Bench

GenEvolve Data and Bench This repository contains the open-source data release for GenEvolve: Config Directory Records Images Purpose…

10K–100K·Apache-2.0·Parquet
ImageMultimodalText

QCalEval

QCalEval Dataset Description The dataset contains scientific plots from quantum computing calibration experiments, paired with vision-language…

<1K·CC-BY·Parquet
ImageMultimodalText

FineBench

Dataset Card for FineBench FineBench is a large-scale, multiple-choice Video Question Answering (VQA) dataset designed specifically to evaluate…

100K–1M·MIT·JSON
MultimodalTextVideo

VideoKR-Eval

VideoKR-Eval 📄 ArXiv  |  💻 Code  |  🤗 Collection About This repository contains the VideoKR-Eval benchmark presented in…

1K–10K·Apache-2.0·JSON
ImageMultimodalTabular

MVBench

MVBench Forked from for reproducibility. Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D…

1K–10K·MIT·JSON
ImageMultimodalText

PhyX

PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…

10K–100K·MIT·Parquet
Text

OpenDocVQA

Dataset Card for OpenDocVQA This is a training and evaluation QA data file for VDocRAG, a new RAG…

10K–100K·Parquet
Advertisement