1K–10K
598 results
MMBench-ru
MMBench-ru This is a translated version of original MMBench dataset and stored in format supported for lmms-eval pipeline.…
UrbanVideo-Bench
ACL'25 Oral UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces This repository contains…
XLRS-Bench-lite
🐙GitHub Information or evaluatation on this dataset can be found in this repo: 📜Dataset License Annotations of this…
CharXiv
CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs NeurIPS 2024 🏠Home (🚧Still in construction) 🤗Data 🥇Leaderboard…
MindCraft
MindCraft Benchmark for spatio-temporal reasoning during vision-and-language navigation, used with LASAR. Layout mindcraft/ 000001/ data.npz 1.mp4…
ANAKIN: manipulated videos and mask annotations
ANAKIN: manipulated videos and mask annotations
FUSION-Finetune-12M
FUSION-12M Dataset Please see paper & website for more information: Overview FUSION-12M is a large-scale, diverse multimodal instruction-tuning…