Skip to content
Advertisement

Image

1,418 results

Image

gtsrb

Dataset Card for German Traffic Sign Recognition Benchmark This dataset contains images of 43 classes of traffic signs.…

100K–1M·Parquet
ImageMultimodalVideo

360Motion-Dataset

360°-Motion Dataset Project page Paper Code Acknowledgments We thank Jinwen Cao, Yisong Guo, Haowen Ji, Jichao Wang, and…

<1K·Apache-2.0·Images (folder)
ImageMultimodalText

fine-t2i

Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning arxiv by Xu Ma, Yitian Zhang, Qihua…

100K–1M·Apache-2.0·WebDataset
Image

OmniDocBench

OmniDocBench English 简体中文 OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, with the following…

1K–10K·Images (folder)
ImageMultimodalText

commoncatalog-cc-by-sa

Dataset Card for CommonCatalog CC-BY-SA This dataset is a large collection of high-resolution Creative Common images (composed of…

1M–10M·CC-BY-SA·Parquet
Image

eurosat

Dataset Card for EuroSAT Dataset Source Paper with code Usage from datasets import load dataset dataset = load…

10K–100K·Parquet
Image

sun397

SUN397 dataset The database contains 397 categories subset from the SUN dataset for Scene Recognition used in the…

10K–100K·Parquet
Image

onepiece

Bangumi Image Base of One Piece This is the image base of bangumi One Piece, we detected 303…

10K–100K·MIT
ImageMultimodalText

document-haystack

Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document…

ImageMultimodalText

kb-books

open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish…

1M–10M·CC0·Arrow
ImageMultimodalText

GeoDrive-Bench

GeoDrive-Bench A multi-country driving scene benchmark for evaluating vision-language models on culture- and region-specific traffic knowledge.…

1K–10K·CC-BY·Parquet
ImageMultimodalText

hle

!NOTE IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing…

1K–10K·MIT·Parquet
ImageMultimodalText

MMMU

This is a merged version of MMMU/MMMU with all subsets concatenated. Large-scale Multi-modality Models Evaluation Suite Accelerating the…

10K–100K·Parquet
Advertisement