Skip to content
Advertisement

Multiple Choice

46 results

ImageMultimodalText

MMMU_Pro

MMMU-Pro (A More Robust Multi-discipline Multimodal Understanding Benchmark) 🌐 Homepage 🏆 Leaderboard 🤗 Dataset 🤗 Paper 📖 arXiv…

1K–10K·Apache-2.0·Parquet
MultimodalTabularText

KMMLU

KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…

100K–1M·Other·CSV
Video

VEU-Bench

Video Editing Understanding(VEU) Benchmark 🖥 Project Page Widely shared videos on the internet are often edited. Recently, although…

Apache-2.0
Text

LongBench-v2

LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: 💻 Github Repo: 📚…

<1K·Apache-2.0·JSON
Text

C-Eval

C-Eval

10K–100K·CC-BY-NC-SA·Parquet
ImageMultimodalText

MMStar

MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper…

1K–10K·Parquet
ImageMultimodalText

GeoDrive-Bench

GeoDrive-Bench A multi-country driving scene benchmark for evaluating vision-language models on culture- and region-specific traffic knowledge.…

1K–10K·CC-BY·Parquet
Advertisement