Skip to content
Advertisement

100K–1M

595 results

ImageMultimodalText

hot3d

HOT3D-Clips This Hugging Face repository hosts HOT3D-Clips, a set of curated sub-sequences of the HOT3D dataset. Download instructions…

100K–1M·WebDataset
Image

gtsrb

Dataset Card for German Traffic Sign Recognition Benchmark This dataset contains images of 43 classes of traffic signs.…

100K–1M·Parquet
ImageMultimodalText

fine-t2i

Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning arxiv by Xu Ma, Yitian Zhang, Qihua…

100K–1M·Apache-2.0·WebDataset
ImageMultimodalSensor / Time-series

libero

This dataset was created using LeRobot. Dataset Description This dataset combines four individual Libero datasets: Libero-Spatial, Libero-Object,…

100K–1M·CC-BY·Parquet
ImageMultimodalText

object365

Objects365 Dataset Objects365 detection dataset in HuggingFace parquet format. Schema Column Type Description image Image RGB image (PIL)…

100K–1M·CC-BY·Parquet
Image

Food-101

Food-101

100K–1M·Custom / Research-only·Parquet
ImageMultimodalText

SpatialEdit-500K

SpatialEdit-500K SpatialEdit-500K is a synthetic training dataset for fine-grained image spatial editing. It is built for learning geometry-aware…

100K–1M·Apache-2.0·WebDataset
ImageMultimodalText

VisionArena-Chat

VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes 200k single and multi-turn chats between users and VLM's…

100K–1M·Parquet
Image

IntuitivePhysics

WorldBench WorldBench is a new benchmark designed to evaluate the physical understanding and prediction of modern world models…

100K–1M·Images (folder)
Image

COCO

330K images with object, segmentation and caption labels.

100K–1M·CC-BY·COCO JSON
Audio

LibriSpeech

~1,000 hours of aligned read English speech.

100K–1M·CC-BY·Audio (wav/mp3)
Advertisement