Multiple Choice
46 results
MMMU_Pro
MMMU-Pro (A More Robust Multi-discipline Multimodal Understanding Benchmark) 🌐 Homepage 🏆 Leaderboard 🤗 Dataset 🤗 Paper 📖 arXiv…
Game of 24 Mathematical Puzzle Dataset
Game of 24 Mathematical Puzzle Dataset
CLUE: Chinese Language Understanding Evaluation benchmark
CLUE: Chinese Language Understanding Evaluation benchmark
KMMLU
KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…
BalkanBench SuperGLUE – Serbian
BalkanBench SuperGLUE - Serbian
MMLU-ProX Multilingual Model Predictions
MMLU-ProX Multilingual Model Predictions
LongBench-v2
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: 💻 Github Repo: 📚…
MMStar
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper…
GeoDrive-Bench
GeoDrive-Bench A multi-country driving scene benchmark for evaluating vision-language models on culture- and region-specific traffic knowledge.…