Apache-2.0
807 results
OpenR1-Math-220k
OpenR1-Math-220k Dataset description OpenR1-Math-220k is a large-scale dataset for mathematical reasoning. It consists of 220k math problems with…
Vchitect_T2V_DataVerse
Vchitect-T2V-Dataverse Vchitect Team1 1Shanghai Artificial Intelligence Laboratory Paper Project Page Data Overview The Vchitect-T2V-Dataverse is the…
LongBench-v2
LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks 🌐 Project Page: 💻 Github Repo: 📚…
AudioMCQ-StrongAC-GeminiCoT
ICLR 2026 DCASE 2026 Training Set AudioMCQ-StrongAC-GeminiCoT This dataset is a highly curated subset of the AudioMCQ dataset,…
ProLongVid_data
Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only…
L2D
TL;DR of L2D, the world's largest self-driving dataset! Read more about L2D on the official Huggingface blog: LeRobot…
OpenThoughts-114k
!NOTE We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-114k Open synthetic reasoning dataset with…
agibot_alpha_v30
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…
VoxSafeBench
VoxSafeBench Demopage: demopage/ Code: This dataset is uploaded as raw files (JSONL + audio), not parquet. Subset: Safety-tier1…