Apache-2.0
807 results
github-code-clean
The GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code…
The dataset of the most popular text-to-image prompts
The dataset of the most popular text-to-image prompts
FineFineWeb
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming…
colored_mnist_28
Colored MNIST Dataset A comprehensive dataset of MNIST digits with RGB colored backgrounds, designed for multi-objective classification tasks…
Vero-600k
Vero-600k Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with…
BLINK
BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage 💻 Code 📖 Paper 📖 arXiv…
FLUX-Reason-6M
FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative…
GSA_volc
GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…
LongVT-Source
LongVT-Source This repository contains the source video and image files for the LongVT project. Overview LongVT is an…
libero
This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "panda", "total episodes":…
360Motion-Dataset
360°-Motion Dataset Project page Paper Code Acknowledgments We thank Jinwen Cao, Yisong Guo, Haowen Ji, Jichao Wang, and…
fine-t2i
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning arxiv by Xu Ma, Yitian Zhang, Qihua…
DocVQA
Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…
PhysicalAI-Robotics-Locomanipulation-GRAIL
📢 News 2026-07-15 Released task-general tracking policy checkpoints trained on the released data. Follow the tracking doc to…