Skip to content
Advertisement

Datasets

1,940 results

ImageMultimodalText

Leopard-Instruct

Leopard-Instruct Paper Github Models-LLaVA Models-Idefics2 Summaries Leopard-Instruct is a large instruction-tuning dataset, comprising 925K…

1M–10M·Apache-2.0·Parquet
ImageMultimodalText

RoadmapBench

RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project…

<1K·MIT·JSON
ImageMultimodalText

the_cauldron

Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a…

1M–10M·Parquet
ImageMultimodalText

FineVision

Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B…

10M–100M·Parquet
3D / Point CloudImageMultimodal

CADS-dataset

CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust,…

10K–100K·Custom / Research-only·CSV
3D / Point CloudImageMultimodal

CADS-dataset

CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust,…

10K–100K·Custom / Research-only·CSV
3D / Point CloudImageMultimodal

CADS-dataset

CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust,…

10K–100K·Custom / Research-only·CSV
ImageMultimodalText

LLaVA-OneVision-2-Data

LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used…

<1K·Apache-2.0·Parquet
Text

The Pile

825 GiB of diverse text from 22 curated sources.

>10M·Mixed·Parquet
Advertisement