OpenAssistant Conversations
OpenAssistant Conversations
1,940 results
OpenAssistant Conversations
Dataset Description You are provided with a large number of Wikipedia comments which have been labeled by human…
PluRel A large collection of 2000 synthetic relational databases (seeds plurel-0 … plurel-1999) generated for pretraining relational/tabular…
Chronos datasets
FineWeb-Edu 100BT (Shuffled)
This is the OSS-Instruct dataset generated by gpt-3.5-turbo-1106 developed by OpenAI. Please pay attention to OpenAI's usage policy…
EuroWeb-2512 EuroWeb is a dataset of collecting multilingual web data from various sources. It was processed with standard…
📐 FineMath What is it? 📐 FineMath consists of 34B tokens (FineMath-3+) and 54B tokens (FineMath-3+ with InfiMM-WebMath-3+)…
ResearchClawBench Evaluating AI Agents for Automated Research from Re-Discovery to New-Discovery Quick Start…
Polymarket Data Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze A comprehensive dataset of 1.9 billion trading…
MolmoBot-data Training episode data (actions, visual inputs, and other sensor data) for 8 tasks on 2 robotic platforms:…
Dataset arXiv Website Code VSI-Bench VSI-Bench-Debiased !IMPORTANT Nov. 7, 2025 UPDATE: This Dataset has been updated to include…
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions M3arsSynth Dataset Summary M3arsSynth is a large-scale,…
BridgeData V2 Video Dataset This dataset contains videos from BridgeData V2 trajectories in VideoFolder format. Derived From This…
Overview RoboCerebra is a large-scale benchmark designed to evaluate robotic manipulation in long-horizon tasks, shifting the focus of…
PhyWorldBench This repository hosts the core assets of PhyWorldBench, the 1,050 JSON prompt files, the evaluation standards, and…
MolmoAct2-DROID Dataset This dataset was created using LeRobot. Language Annotations This dataset includes annotated language instructions in…
REVISOR-25k A multi-task video understanding dataset for training video LLMs with reinforcement learning (GRPO). The dataset contains ~25k…