StockChina-Minute
A-Share Minute-Level Historical Data Dataset Description This dataset contains minute-level trading data for Chinese A-share stocks from 2005…
2,368 results
A-Share Minute-Level Historical Data Dataset Description This dataset contains minute-level trading data for Chinese A-share stocks from 2005…
XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar)…
osm-polygon-selection dataset A curated set of OpenStreetMap polygons from 310 geographic units — sovereign countries plus sub-country regions…
Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from…
ManaTTS-Persian-Speech-Dataset ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality…
This dataset contains the fully prepared data, which has been tokenized and pre-shuffled, used to train the Pythia…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "frodobot", "total episodes":…
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper…
Core-S2L1C Contains a global coverage of Sentinel-2 (Level 1C) patches, each of size 1,068 x 1,068 pixels. Source…
This repo contains (up to) 30k samples of 21 languages (top 20 languages by StackOverflow survey, html/css was…
Romanian Real Estate Listings Dataset
Hate Speech and Offensive Language
Code Leaderboard Results Paper RewardBench 2 Evaluation Dataset Card The RewardBench 2 evaluation dataset is the new version…
MVTamperBench Dataset Overview MVTamperBench is a robust benchmark designed to evaluate Vision-Language Models (VLMs) against adversarial video…
Languages: 简体中文 · English dojo forex kline — FX Daily Bars Overview Daily OHLC and amplitude for major…
Dataset Summary This dataset is a modified version of the Emilia corpus, converted into parquet format to facilitate…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v2.1", "robot type": "dual panda robotiq…
AIDev: Studying AI Coding Agents on GitHub (The Rise of AI Teammates in Software Engineering 3.0) Papers: The…
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v2.0", "robot type": "franka", "total episodes":…
Pretraining (MimicGen) — atomic MimicGen-generated rollouts across 60 atomic tasks (~10,000 demos/task). 1,615 hours total, generated by scripted…