Skip to content
Advertisement

Text-to-Image

52 results

ImageMultimodalText

imagenet-1k-vl-enriched

Visualize on Visual Layer Imagenet-1K-VL-Enriched An enriched version of the ImageNet-1K Dataset with image caption, bounding boxes, and…

1M–10M·Apache-2.0·Parquet
3D / Point Cloud

Objaverse-Ortho10View

Objaverse-Ortho10View Github Project Page Paper 1. Dataset Introduction TL;DR: This dataset contains multi-view images that are rendered from…

MIT
3D / Point Cloud

Objaverse-Rand6View

Objaverse-Rand6View Github Project Page Paper 1. Dataset Introduction TL;DR: This dataset contains multi-view images that are rendered from…

MIT
MultimodalTextVideo

Vript

🎬 Vript: Refine Video Captioning into Video Scripting Github Repo We construct a fine-grained video-text dataset with 12K…

100K–1M·JSON
ImageMultimodalTabular

Recap-DataComp-1B

Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced…

>1B·CC-BY·Parquet
ImageMultimodalText

commoncatalog-cc-by

Dataset Card for CommonCatalog CC-BY This dataset is a large collection of high-resolution Creative Common images (composed of…

10M–100M·CC-BY·Parquet
ImageMultimodalText

commoncatalog-cc-by-sa

Dataset Card for CommonCatalog CC-BY-SA This dataset is a large collection of high-resolution Creative Common images (composed of…

1M–10M·CC-BY-SA·Parquet
Advertisement