WhestBench 2026: ARC White-Box Estimation Challenge
WhestBench 2026: ARC White-Box Estimation Challenge
2,368 results
WhestBench 2026: ARC White-Box Estimation Challenge
Dataset Card for "QM9" QM9 dataset from Ruddigkeit et al., 2012; Ramakrishnan et al., 2014. Original data downloaded…
cc100-documents This dataset is a restructured version of the CC-100 (statmt/cc100) dataset. In the original dataset, each instance…
Crystallography Open Database (COD) — Full Snapshot A complete mirror of the Crystallography Open Database (COD) as a…
Liars' Bench Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here.…
Dataset Card for CodeForces-CoTs Dataset description CodeForces-CoTs is a large-scale dataset for training reasoning models on competitive…
GroMo 25 — Multiview Plant Growth Dataset
A version of the PersonaChat dataset that has been true-cased, and also has been given more normalized punctuation.…
Dataset for SylReg This repository contains the datasets and alignments associated with the paper Speaker-Disentangled Chunk-Wise Regression for…
Polymarket Crypto Up/Down Markets
Agent Apprenticeship Seed Dataset v0.2
NASA SMAP and MSL Spacecraft Anomaly Detection Dataset