Skip to content
Advertisement

10M–100M

169 results

MultimodalTabularText

agibot_alpha_v30

This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase version": "v3.0", "robot type": "AgiBot A2D", "total…

10M–100M·Apache-2.0·Parquet
Text

github-code-clean

The GitHub Code clean dataset in a more filtered version of codeparrot/github-code dataset, it consists of 115M code…

10M–100M·Apache-2.0
Text

wikipedia

Dataset Card for Wikimedia Wikipedia Dataset Summary Wikipedia dataset containing cleaned articles of all languages. The dataset is…

10M–100M·Mixed·Parquet
ImageMultimodalText

Imagenet21K

NOTE: I have recaptioned all images here This dataset is the entire 21K ImageNet dataset with about 13…

10M–100M·Parquet
ImageMultimodalText

chitralekha

Chitralekha Dataset Details Dataset Version Some of the fonts do not have proper letters/rendering of different telugu letter…

10M–100M·MIT·Parquet
ImageMultimodalText

commoncatalog-cc-by

Dataset Card for CommonCatalog CC-BY This dataset is a large collection of high-resolution Creative Common images (composed of…

10M–100M·CC-BY·Parquet
ImageMultimodalText

GQA

Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…

10M–100M·MIT·Parquet
ImageMultimodalText

latex-formulas-80M

For more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset…

10M–100M·Apache-2.0·Parquet
ImageMultimodalText

cc12m-wds

Dataset Card for Conceptual Captions 12M (CC12M) Dataset Summary Conceptual 12M (CC12M) is a dataset with 12 million…

10M–100M·Custom / Research-only·WebDataset
ImageMultimodalText

RenderedText

This dataset has been created by Stability AI and LAION. This dataset contains 12 million 1024x1024 images of…

10M–100M·WebDataset
ImageMultimodalTabular

MapPool

MapPool - Bubbling up an extremely large corpus of maps for AI MapPool is a dataset of 75…

10M–100M·CC-BY·Parquet
ImageMultimodalText

FineVision

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B…

10M–100M·Parquet
Advertisement