KodCode-V1
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset…
2,175 results
🐱 KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding KodCode is the largest fully-synthetic open-source dataset…
Languages: 简体中文 · English dojo quote — Latest Quote Snapshot Overview Point-in-time snapshot for the most recent trading…
UCI Human Activity Recognition (HAR) Dataset Dataset Description The UCI Human Activity Recognition dataset is a widely-used benchmark…
India Index & Options 1-minute OHLC
Paired Llama 3.2 1B Token Embeddings (LMSYS-Chat-1M) This dataset contains paired activations corresponding to single token locations extracted…
Overview Property Value Source Project Gutenberg (English catalog) Snapshot 2026-07-02-18-47-04 Total files 50871 Total Tokens (BPE) ~7.14 billion…
Paper You can find more information in our paper. Dataset Paper:
BTCUSDT Perpetual Futures — 5-Minute Feature Dataset Complete historical dataset for Binance BTCUSDT USDT-Margined Perpetual Futures, covering…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "fps": 20, "features": {…
CodeParrot 🦜 Dataset Cleaned (valid) Train split of CodeParrot 🦜 Dataset Cleaned. Dataset structure DatasetDict({ train: Dataset({ features:…
Dataset Information This dataset includes daily price data for various stocks. Instruments Included 7000+ US Stocks Dataset Columns…
KLING AI Generative Media Dataset
The Heap Dataset We develop The Heap, a new contamination-free multilingual code dataset comprising 57 languages, which facilitates…
BeyondArena Datasets