SWE-bench_Pro
Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…
3,216 results
Dataset Summary SWE-Bench Pro is a challenging, enterprise-level dataset for testing agent ability on long-horizon software engineering tasks.…
Common Corpus Full paper - ICLR 2026 oral Common Corpus is the largest open licensed text dataset, comprising…
Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The…
Wikitext Document Level This is a modified version of that returns Wiki pages instead of Wiki text line-by-line.…
Shofo TikTok General (Small) Overview Shofo TikTok General (Small) is a dataset containing 50,000 TikTok videos with comprehensive…
!NOTE We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-114k Open synthetic reasoning dataset with…
🚀 AutoMathText-V2: A 2.46 Trillion Token AI-Curated STEM Pretraining Dataset 🎉 AutoMathText-v2 has surpassed 1.5 million downloads!…
Dataset Summary SWE-bench Verified is a subset of 500 samples from the SWE-bench test set, which have been…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…
The Cross-lingual TRansfer Evaluation of Multilingual Encoders for Speech (XTREME-S) benchmark is a benchmark designed to evaluate speech…
Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in…
Annotation We resized the dataset to 1080p for easier uploading. Therefore, the original annotation file might not match…