Skip to content
Advertisement

Apache-2.0

807 results

Text

xP3x

xP3x

100M–1B·Apache-2.0·Parquet
Text

github-code-clean

The GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code…

10M–100M·Apache-2.0
MultimodalTabularText

FineFineWeb

FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming…

>1B·Apache-2.0
Image

colored_mnist_28

Colored MNIST Dataset A comprehensive dataset of MNIST digits with RGB colored backgrounds, designed for multi-objective classification tasks…

10K–100K·Apache-2.0·Images (folder)
ImageMultimodalText

Vero-600k

Vero-600k Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with…

100K–1M·Apache-2.0·Parquet
ImageMultimodalText

BLINK

BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage 💻 Code 📖 Paper 📖 arXiv…

1K–10K·Apache-2.0·Parquet
ImageMultimodalText

FLUX-Reason-6M

FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative…

1M–10M·Apache-2.0·Parquet
ImageMultimodalText

GSA_volc

GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…

1M–10M·Apache-2.0·JSON
Image

FinErva

🌌 FinErva: Interpretable Multimodal Reasoning for Robo-Advisory A dataset & lightweight training framework that teaches small models to…

1K–10K·Apache-2.0·Images (folder)
ImageMultimodalVideo

LongVT-Source

LongVT-Source This repository contains the source video and image files for the LongVT project. Overview LongVT is an…

<1K·Apache-2.0
ImageMultimodalSensor / Time-series

libero

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "panda", "total episodes":…

100K–1M·Apache-2.0·Parquet
ImageMultimodalVideo

360Motion-Dataset

360°-Motion Dataset Project page Paper Code Acknowledgments We thank Jinwen Cao, Yisong Guo, Haowen Ji, Jichao Wang, and…

<1K·Apache-2.0·Images (folder)
ImageMultimodalText

fine-t2i

Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning arxiv by Xu Ma, Yitian Zhang, Qihua…

100K–1M·Apache-2.0·WebDataset
ImageMultimodalText

DocVQA

Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…

10K–100K·Apache-2.0·Parquet
Advertisement