Skip to content
Advertisement

Image Text To Text

58 results

ImageMultimodalText

MGRS-200k

MGRS-200k Dataset MGRS-200k is the first multi-granularity remote sensing (RS) image-text dataset, introduced in the paper FarSLIP: Discovering…

100K–1M·WebDataset
Image

sentinel-lfm-mining-patches

sentinel-lfm — illegal-mining single-frame patches 128px RGB patches cropped from the Roboflow illegal-mining dataset, labelled mine (1) /…

1K–10K·Custom / Research-only·Images (folder)
AudioImageMultimodal

SoundingEarth

SoundingEarth SoundingEarth is a geo-referenced soundscape dataset that pairs Google Earth imagery with geotagged environmental audio recordings…

10K–100K·CC-BY·Parquet
ImageMultimodalText

DocVQA-2026

DocVQA 2026 ICDAR2026 Competition on Multimodal Reasoning over Documents in Multiple Domains Building upon previous DocVQA benchmarks, this…

<1K·Parquet
ImageMultimodalText

SDG-30K

SDG-30K — Structured Defect Grounding Dataset A 30,000-image dataset for structured defect grounding in text-to-image generations. Each image…

10K–100K·CC-BY-NC·JSON
Image

PanoEnv

CVPR 2026 Highlight - PanoEnv-QA: A Large-Scale Geometry-Grounded Panoramic VQA Benchmark for 3D Spatial Intelligence 📖 Overview PanoEnv-QA…

10K–100K·CC-BY·Images (folder)
ImageMultimodalText

ChartVerse-SFT-600K

ChartVerse-SFT-600K is a large-scale, high-quality chart reasoning dataset with Chain-of-Thought (CoT) annotations, developed as part of the…

100K–1M·Apache-2.0·Parquet
Advertisement