Skip to content
Advertisement

Text

2,175 results

Text

github-code-2025-language-split

📜 Source Data & Attribution This dataset is a processed derivative of nick007x/github-code-2025. Origination The original data was…

100M–1B·Custom / Research-only·Parquet
Text

OpenThoughts-1k-sample

!NOTE We have released a paper for OpenThoughts! See our paper here. Open-Thoughts-1k-sample This is a 1k sample…

1K–10K·Parquet
ImageMultimodalText

COCO-Caption

Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…

10K–100K·Parquet
ImageMultimodalText

Imagenet21K

NOTE: I have recaptioned all images here This dataset is the entire 21K ImageNet dataset with about 13…

10M–100M·Parquet
ImageMultimodalTabular

FIP1

The FIP 1.0 Data Set: Highly Resolved Annotated Image Time Series of 4,000 Wheat Plots Grown in Six…

1K–10K·CC-BY·Parquet
ImageMultimodalText

ai2d

@misc{kembhavi2016diagram, title={A Diagram Is Worth A Dozen Images}, author={Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Minjoon…

1K–10K·Parquet
ImageMultimodalText

Vero-600k

Vero-600k Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with…

100K–1M·Apache-2.0·Parquet
ImageMultimodalText

BLINK

BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage 💻 Code 📖 Paper 📖 arXiv…

1K–10K·Apache-2.0·Parquet
ImageMultimodalTabular

Recap-DataComp-1B

Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced…

>1B·CC-BY·Parquet
ImageMultimodalText

megalith-mdqa

Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.

1M–10M·OpenRAIL·Parquet
ImageMultimodalText

FLUX-Reason-6M

FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative…

1M–10M·Apache-2.0·Parquet
ImageMultimodalText

GSA_volc

GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…

1M–10M·Apache-2.0·JSON
ImageMultimodalText

MMStar

MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper…

1K–10K·Parquet
Advertisement