Skip to content
Advertisement

1M–10M

329 results

MultimodalTextVideo

ProLongVid_data

Dataset Card for ProLongVid-data Uses This dataset is used for the training of the ProLongVid model. We only…

1M–10M·Apache-2.0·JSON
Text

TinyStories

Dataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in…

1M–10M·Other·Parquet
Text

wikipedia-monthly

🚀 Wikipedia Monthly Last updated: March 14, 2026, 21:06 UTC This repository provides monthly, multilingual dumps of Wikipedia,…

1M–10M·CC-BY-SA
ImageMultimodalVideo

GOKU-2M

Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing GOKU-2M is a large-scale, unified instruction-based…

1M–10M·CC-BY-NC
ImageMultimodalText

megalith-mdqa

Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.

1M–10M·OpenRAIL·Parquet
ImageMultimodalText

FLUX-Reason-6M

FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative…

1M–10M·Apache-2.0·Parquet
ImageMultimodalText

GSA_volc

GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…

1M–10M·Apache-2.0·JSON
Image

yanto991

GroundCUA: Grounding Computer Use Agents on Human Demonstrations 🌐 Website 📑 Paper 🤗 Dataset 🤖 Models GroundCUA Dataset…

1M–10M·MIT
ImageMultimodalText

idl-wds

Dataset Card for Industry Documents Library (IDL) Dataset Summary Industry Documents Library (IDL) is a document dataset filtered…

1M–10M·Custom / Research-only·WebDataset
ImageMultimodalText

ChartNet

ChartNet: A Million-Scale Multimodal Dataset for Chart Understanding 🌐 Homepage 📖 arXiv 📝 Changelog June 3, 2026 —…

1M–10M·Other·Parquet
Image

ImageNet

ImageNet

1M–10M·Custom / Research-only·Parquet
ImageMultimodalText

cc3m-wds

Dataset Card for Conceptual Captions (CC3M) Dataset Summary Conceptual Captions is a dataset consisting of ~3.3M images annotated…

1M–10M·Custom / Research-only·WebDataset
ImageMultimodalText

commoncatalog-cc-by-sa

Dataset Card for CommonCatalog CC-BY-SA This dataset is a large collection of high-resolution Creative Common images (composed of…

1M–10M·CC-BY-SA·Parquet
Advertisement