Skip to content
Advertisement

Multimodal

2,368 results

ImageMultimodalVideo

360Motion-Dataset

360°-Motion Dataset Project page Paper Code Acknowledgments We thank Jinwen Cao, Yisong Guo, Haowen Ji, Jichao Wang, and…

<1K·Apache-2.0·Images (folder)
ImageMultimodalText

fine-t2i

Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning arxiv by Xu Ma, Yitian Zhang, Qihua…

100K–1M·Apache-2.0·WebDataset
ImageMultimodalText

commoncatalog-cc-by-sa

Dataset Card for CommonCatalog CC-BY-SA This dataset is a large collection of high-resolution Creative Common images (composed of…

1M–10M·CC-BY-SA·Parquet
ImageMultimodalText

document-haystack

Document Haystack Dataset This repository contains the dataset for the paper “Document Haystack: A Long Context Multimodal Image/Document…

ImageMultimodalText

kb-books

open-rdl-books Dataset Description Language dan, dansk, Danish License Public Domain, cc0-1.0 Dataset Summary Documents from the Royal Danish…

1M–10M·CC0·Arrow
ImageMultimodalText

GeoDrive-Bench

GeoDrive-Bench A multi-country driving scene benchmark for evaluating vision-language models on culture- and region-specific traffic knowledge.…

1K–10K·CC-BY·Parquet
ImageMultimodalText

hle

!NOTE IMPORTANT: Please help us protect the integrity of this benchmark by not publicly sharing, re-uploading, or distributing…

1K–10K·MIT·Parquet
ImageMultimodalText

MMMU

This is a merged version of MMMU/MMMU with all subsets concatenated. Large-scale Multi-modality Models Evaluation Suite Accelerating the…

10K–100K·Parquet
ImageMultimodalText

DocVQA

Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…

10K–100K·Apache-2.0·Parquet
ImageMultimodalText

OpenFake

Dataset Card for OpenFake OpenFake is a dataset and benchmark for detecting AI-generated images, with a focus on…

1M–10M·CC-BY-NC·Parquet
ImageMultimodalTabular

MVBench

MVBench Important Update 18/10/2024 Due to NTU RGB+D License, 320 videos from NTU RGB+D need to be downloaded…

1K–10K·MIT·JSON
ImageMultimodalText

textvqa

Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…

10K–100K·Parquet
ImageMultimodalText

GQA

Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…

10M–100M·MIT·Parquet
Advertisement