Skip to content
Advertisement
MultimodalTabularText

MMLU_PromptEval_full

MMLU PromptEval full

MMLU Multi-Prompt Evaluation Data Overview This dataset contains the results of a comprehensive evaluation of various Large Language Models (LLMs) using multiple prompt templates on the Massive Multitask Language Understanding (MMLU) benchmark. The data is introduced in Maia Polo, Felipe, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin. “Efficient multi-prompt evaluation of LLMs.”… See the full description on the dataset page: MMLU full.

Source: Hugging Face Hub (PromptEval/PromptEval_MMLU_full). Metadata imported from the dataset’s Hub tags.

Advertisement