Tabular
1,117 results
Creative Professionals Agentic Tasks (1M)
Creative Professionals Agentic Tasks (1M)
daVinci-Dev (Agent-native Trajectories)
daVinci-Dev (Agent-native Trajectories)
robocasa_merged_24_tasks_100demos_v1
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v2.1", "robot type": "panda", "total episodes":…
dementor-matrix-responses
Dementor — matrix model responses Generated model outputs for the Dementor LLM-imitation / behavioral-inertia study. Companion to: Code…
The Vuk’uzenzele South African Multilingual Corpus
The Vuk'uzenzele South African Multilingual Corpus
MSR-VTT
Clone from "friedrichor/MSR-VTT". MSRVTT contains 10K video clips and 200K captions. We adopt the standard 1K-A split protocol,…
X-Atlas-Orion
X-Atlas/Orion X-Atlas: Orion edition (X-Atlas/Orion) is a Perturb-seq atlas containing two genome-wide Fix-Cryopreserve-ScRNAseq (FiCS) Perturb-seq…
fishbase
Dataset Card for FishBase Snapshots of FishBase data tables used by the rOpenSci package rfishbase and the FishBase…
FlyRank Internship — Warehouse Star Schema (Pseudonymized, Gated)
FlyRank Internship — Warehouse Star Schema (Pseudonymized, Gated)
StockChina-Minute
A-Share Minute-Level Historical Data Dataset Description This dataset contains minute-level trading data for Chinese A-share stocks from 2005…
xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025 This dataset contains historical price data for XAUUSD (Gold vs US Dollar)…
osm-polygon-selection
osm-polygon-selection dataset A curated set of OpenStreetMap polygons from 310 geographic units — sovereign countries plus sub-country regions…
AIME_1983_2024
Disclaimer: This is a Benchmark dataset! Do not using in training! This is the Benchmark of AIME from…
Mana-TTS
ManaTTS-Persian-Speech-Dataset ManaTTS is the largest publicly available single-speaker Persian corpus, comprising over 114 hours of high-quality…
pile-deduped-pythia-preshuffled
This dataset contains the fully prepared data, which has been tokenized and pre-shuffled, used to train the Pythia…
berkeley-frodobots-lerobot-7k
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "frodobot", "total episodes":…
medicine-tasks
Adapting LLMs to Domains via Continual Pre-Training (ICLR 2024) This repo contains the evaluation datasets for our paper…