AIME_1983_2024
Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from…
173 results
Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from…
Fajri Koto, Haonan Li, Sara Shatnawi, Jad Doughman, Abdelrahman Boda Sadallah, Aisha Alraeesi, Khalid Almubarak, Zaid Alyafeai, Neha…
Artificial Analysis Long Context Reasoning (AA-LCR) Dataset AA-LCR includes 100 hard text-based questions that require reasoning across multiple…
H&M Personalized Fashion Recommendations
This repository contains the extracted time series features (tsfeatures) for each variate and the detailed forecasting results for…
Dataset Card for Demo1 Dataset Summary This is a demo dataset. It consists in two files data/train.csv and…
GroMo 25 — Multiview Plant Growth Dataset
KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…
This dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about…
PersonaMem v2, Implicit Persona, LLM Personalization
Update 01/31/2024 We update the OpenAI Moderation API results for ToxicChat (0124) based on their updated moderation model…
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust,…
Dataset Card for PopQA Dataset Summary PopQA is a large-scale open-domain question answering (QA) dataset, consisting of 14k…