Skip to content
Advertisement

General

2,594 results

ImageMultimodalText

RoadmapBench

RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project…

<1K·MIT·JSON
Image

Cifar10

Cifar10

10K–100K·Custom / Research-only·Parquet
ImageMultimodalText

the_cauldron

Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a…

1M–10M·Parquet
Image

IntuitivePhysics

WorldBench WorldBench is a new benchmark designed to evaluate the physical understanding and prediction of modern world models…

100K–1M·Images (folder)
Image

badges

Badges A set of badges you can use anywhere. Just update the anchor URL to point to the…

<1K·MIT·Images (folder)
ImageMultimodalText

FineVision

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B…

10M–100M·Parquet
Image

documentation-images

This dataset contains images used in the documentation of HuggingFace's libraries. HF Team: Please make sure you optimize…

<1K·CC-BY-NC-SA·Images (folder)
Image

banned-historical-archives

和谐历史档案馆数据集 - Banned Historical Archives Datasets 和谐历史档案馆数据集包含已录入 和暂未未录入的原始文件。 目录结构 banned-historical-archives.github.io 已录入该网站的原始数据,不定期从 github 仓库中同步…

<1K·Images (folder)
ImageMultimodalText

LLaVA-OneVision-2-Data

LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used…

<1K·Apache-2.0·Parquet
AudioMultimodalVideo

dataset

YouTube Video Downloads This dataset contains locally downloaded YouTube video files from /data/youtube downloads. Source files scanned: 88549…

Custom / Research-only
Video

Dexora_Real-World_Dataset

Dexora: Open-Source VLA for High-DoF Bimanual Dexterity 🔥 News & Updates 2025-12-03: Released the full Real-World Dataset (12.2K…

MIT
Image

COCO

330K images with object, segmentation and caption labels.

100K–1M·CC-BY·COCO JSON
Image

ImageNet

14M+ images across 20K+ categories (1K subset common).

>10M·Custom / Research-only·Images (folder)
Audio

AudioSet

2M+ human-labeled 10s clips across 600+ classes.

1M–10M·CC-BY·Audio (wav/mp3)
Text

The Pile

825 GiB of diverse text from 22 curated sources.

>10M·Mixed·Parquet
Audio

LibriSpeech

~1,000 hours of aligned read English speech.

100K–1M·CC-BY·Audio (wav/mp3)
Advertisement