Skip to content
Advertisement

Images (folder)

369 results

Image

PanoEnv

CVPR 2026 Highlight - PanoEnv-QA: A Large-Scale Geometry-Grounded Panoramic VQA Benchmark for 3D Spatial Intelligence 📖 Overview PanoEnv-QA…

10K–100K·CC-BY·Images (folder)
Image

SandThink

SandThink Dataset (v1.0) SandThink 是一个专为具身智能 (Embodied AI) 任务设计的大规模指令微调与偏好对齐数据集。该数据集通过结构化的 Chain-of-Thought (CoT) 推理过程,显著提升了 Vision-Language-Action…

<1K·MIT·Images (folder)
ImageMultimodalText

HoloCount

HoloCount: A Holistic Visual Counting Benchmark for MLLMs Abstract Visual counting is a fundamental pillar of multimodal intelligence,…

1K–10K·CC-BY·Images (folder)
Image

causalphys

Causal-VL Dataset Causal reasoning VQA dataset with 4 categories × 4 subcategories (3062 questions). Structure Each subcategory contains:…

<1K·Apache-2.0·Images (folder)
Image

C3B

English 简体中文 C³B: Comics Cross-Cultural Benchmark Culture In a Frame: C³B as a Comic-Based Benchmark for Multimodal Cultural…

1K–10K·CC-BY·Images (folder)
ImageMultimodalText

FlowLearn

cr--- task categories: - visual-question-answering language: - en size categories: - 1K<n<10K References The original articles are maintained…

<1K·CC-BY-NC-SA·Images (folder)
ImageMultimodalVideo

MindCraft

MindCraft Benchmark for spatio-temporal reasoning during vision-and-language navigation, used with LASAR. Layout mindcraft/ 000001/ data.npz 1.mp4…

1K–10K·Apache-2.0·Images (folder)
Advertisement