ANAKIN: manipulated videos and mask annotations
ANAKIN: manipulated videos and mask annotations
2,147 results
ANAKIN: manipulated videos and mask annotations
FUSION-12M Dataset Please see paper & website for more information: Overview FUSION-12M is a large-scale, diverse multimodal instruction-tuning…
Molmo2-SynMultiImageQA Molmo2-SynMultiImageQA is a collection of synthetic multi-image question-answer pairs about various kinds of text-rich images,…
MMFineReason-Full-2.3M The Complete Pre-Selection Dataset — Before Quality Filtering 📖 Overview MMFineReason-Full-2.3M is the complete pre-selection…
ULVR-filtered Filtered subset of RuoliuYang/ULVR v2 clean: the 101,951 training samples that Qwen2.5-VL-7B-Instruct answered incorrectly given only…
REVERSE Dataset Dataset for REVERSE (Reinforcing Evidence Verification and Search for Agentic Image Geolocation). This dataset supports training…
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation GUI Unbiasing via…
Cambrian Vision-Centric Benchmark (CV-Bench)
Kvasir-VQA-x1 A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Kvasir-VQA-x1 on GitHub Original Image…
Banana-Merged: Multi-page VQA with Hard Negatives
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation 🤗 Dataset 🖥️ Code 📄 Paper 🔥 News 2026.02…
RoboBench: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models as Embodied Brain 📋 Overview RoboBench is a…
Vietnamese Medicinal Herb VQA