KMMLU
KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…
3,216 results
KMMLU (Korean-MMLU) We propose KMMLU, a new Korean benchmark with 35,030 expert-level multiple-choice questions across 45 subjects ranging…
FAISS Indices and Chunked Datasets for SmolLM and SmolLM2 corpora This repository contains part of the FAISS indices…
ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠…
Dataset Card GPT-NL Public Corpus The GPT-NL Public Corpus is the largest permissively licensed Dutch-language resource available for…
RAGBench Dataset Overview RAGBEnch is a large-scale RAG benchmark dataset of 100k RAG examples. It covers five unique…
Core-S2L2A Contains a global coverage of Sentinel-2 (Level 2A) patches, each of size 1,068 x 1,068 pixels. Source…
TODO: Add YAML tags here. Copy-paste the tags obtained with the online tagging app: annotations creators: - no-annotation…
Filtered WIT, an Image-Text Dataset. A reliable Dataset to run Image-Text models. You can find WIT, Wikipedia Image…
This dataset contains the subset of ArXiv papers with the "cs.LG" tag to indicate the paper is about…
Indoor Safety Hazard Detection & Work-Zone Monitoring Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from…
The Stack v2 Subset with File Contents (Python, Java, JavaScript, C, C++) TempestTeam/dataset-the-stack-v2-dedup-sub Dataset Summary This dataset is…
PersonaMem v2, Implicit Persona, LLM Personalization
Dataset Card for WildGuardMix Disclaimer: The data includes examples that might be disturbing, harmful or upsetting. It includes…
coding agent traces - security audits
📚 Institutional Books 1.0 Institutional Books is a growing corpus of public domain books. This 1.0 release is…
BalkanBench SuperGLUE - Serbian
Ropedia Xperience-10M Task Suite Artifacts