VisionArena-Chat
VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes 200k single and multi-turn chats between users and VLM's…
2,368 results
VisionArena-Battle: 30K Real-World Image Conversations with Pairwise Preference Votes 200k single and multi-turn chats between users and VLM's…
PhysicalAI-WorldModel-Synthetic-Physical-Interaction-Scenes Dataset Card Dataset Description…
personalized visual instruction tuning
MolmoAct - Midtraining Mixture Data Mixture used for MolmoAct Midtraining. Contains MolmoAct Dataset formulated as Action Reasoning Data.…
JL1 CUP 2024 (Second Track — SCD folder layout)
Leopard-Instruct Paper Github Models-LLaVA Models-Idefics2 Summaries Leopard-Instruct is a large instruction-tuning dataset, comprising 925K…
RoadmapBench A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project…
Dataset Card for The Cauldron Dataset description The Cauldron is part of the Idefics2 release. It is a…
LLaVA-OneVision-1.5 Instruction Data Paper Code 📌 Introduction This dataset, LLaVA-OneVision-1.5-Instruct, was collected and integrated during the…
Fine Vision FineVision is a massive collection of datasets with 17.3M images, 24.3M samples, 88.9M turns, and 9.5B…
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust,…
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust,…
CADS: A Comprehensive Anatomical Dataset and Segmentation for Whole-Body Anatomy in Computed Tomography Overview CADS is a robust,…
Unity + SoundSpaces 2.0 Spatial Audio Dataset (Replica)
LLaVA-OneVision-2-Data Training data for the LLaVA-OneVision-2 multimodal model family, covering large-scale video and spatial reasoning corpora used…