Skip to content
Advertisement
MultimodalTabularText

MMLU-ProX Multilingual Model Predictions

MMLU-ProX Multilingual Model Predictions

MMLU-ProX Multilingual Model Predictions Raw per-sample model predictions on MMLU-ProX across 29 languages and 25 open-weight LLMs, produced with lm-evaluation-harness. This dataset releases the full prediction logs (not just aggregate scores) so that item-level responses can be re-analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per-item difficulty estimation. Repository structure mmlu prox / └──… See the full description on the dataset page:

Source: Hugging Face Hub (gililior/mmlu-prox-eval-predictions). Metadata imported from the dataset’s Hub tags.

Advertisement