MapEval-API
MapEval-API MapEval-API is created using MapQaTor. Usage from datasets import load dataset Load dataset ds = load dataset("MapEval/MapEval-API",…
46 results
MapEval-API MapEval-API is created using MapQaTor. Usage from datasets import load dataset Load dataset ds = load dataset("MapEval/MapEval-API",…
MapEval-Visual This dataset was introduced in MapEval: A Map-Based Evaluation of Geo-Spatial Reasoning in Foundation Models Example Query…
CHOICE: Benchmarking The Remote Sensing Capabilities of Large Vision-Language Models Abstract: The rapid advancement of Large Vision-Language Models…
Remote Sensing VQA — Multilingual A multilingual counterfactual MCQ dataset built from remote sensing / satellite imagery. Each…
PhyX: Does Your Model Have the "Wits" for Physical Reasoning? Dataset for the paper "PhyX: Does Your Model…
AVQA (Audio-Visual QA) — Videos + Annotations
Dataset Card for MathVerse Dataset Description Paper Information Dataset Examples Leaderboard Citation Dataset Description The capabilities of…
Dataset Card for MathVerse This is the version for lmms-eval. This shares the same data with the official…
ShipBench: A Drawing-Grounded VLM Benchmark for Ship Structural Reasoning ShipBench is a metadata-grounded vision-language benchmark on…
🐙GitHub Information or evaluatation on this dataset can be found in this repo: 📜Dataset License Annotations of this…