COCO-Caption
Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…
1,418 results
Large-scale Multi-modality Models Evaluation Suite Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval 🏠 Homepage…
NOTE: I have recaptioned all images here This dataset is the entire 21K ImageNet dataset with about 13…
Bangumi Image Base of A-rank Party Wo Ridatsu Shita Ore Wa, Moto Oshiego-tachi To Meikyuu Shinbu Wo Mezasu.…
Negative Embedding This is a Negative Embedding trained with Counterfeit. Please use it in the "stable-diffusion-webuiembeddings" folder.It can…
Colored MNIST Dataset A comprehensive dataset of MNIST digits with RGB colored backgrounds, designed for multi-objective classification tasks…
0426 CRef/SRef LoRA Triplet Dataset
The FIP 1.0 Data Set: Highly Resolved Annotated Image Time Series of 4,000 Wheat Plots Grown in Six…
Goku: A Million-Scale Universal Dataset and Benchmark for Instruction-Based Video Editing GOKU-2M is a large-scale, unified instruction-based…
@misc{kembhavi2016diagram, title={A Diagram Is Worth A Dozen Images}, author={Aniruddha Kembhavi and Mike Salvato and Eric Kolve and Minjoon…
ASEAN Geolocation
Vero-600k Vero is a fully open reinforcement learning (RL) recipe for training and evaluating multi-task visual reasoning with…
Synthetic Veterinary Ultrasound Dataset (AFAST)
Game Character Skins Dataset Summary This comprehensive dataset contains game character skins and artwork from multiple popular mobile…
KITTI Pseudo Depth (Eigen Split) with Depth Anything V2 Dataset Description This dataset provides high-quality pseudo depth maps…
BLINK: Multimodal Large Language Models Can See but Not Perceive 🌐 Homepage 💻 Code 📖 Paper 📖 arXiv…