JSON
248 results
Nemotron-ClimbMix
ClimbMix Dataset 🚀 Creating the highest-quality pre-training datasets for LLMs 🌟 📄 PAPER 🤗 CLIMBLAB 🤗 CLIMBMIX 🏠…
coding agent traces – security audits
coding agent traces - security audits
AlgoTune
Website Paper Code How good are language models at coming up with new algorithms? To try to answer…
Finch
Finch (FinWorkBench): Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows This repository contains the dataset for…
FineFineWeb-sample
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming…
hhtools_parc_ms
hhtools PARC MS — terrain-aware humanoid motion clips 中文说明 Motion clips in PARC MS layout for human-humanoid-tools (hhtools):…
MMLU-ProX Multilingual Model Predictions
MMLU-ProX Multilingual Model Predictions
EU Law Dataset – Category 15.10
EU Law Dataset - Category 15.10
topological-traps-dataset
Topological Traps Dataset TAU Algorithmic Robotics - Fall 2025/2026 - Daniel Simanovsky Pre-processed occupancy grids and Oracle viability…
PKU-SafeRLHF
Dataset Card for PKU-SafeRLHF Warning: this dataset contains data that may be offensive or harmful. The data are…
the-stack-smol
Dataset Description A small subset (~0.1%) of the-stack dataset, each programming language has 10,000 random samples from the…
car-bench-dataset
CAR-Bench Dataset CAR-Bench is a benchmark for evaluating AI voice assistants in a realistic automotive (car) environment. It…
Magicoder-OSS-Instruct-75K
This is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy…
ResearchClawBench
ResearchClawBench Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery Quick Start…