MGRS-200k
MGRS-200k Dataset MGRS-200k is the first multi-granularity remote sensing (RS) image-text dataset, introduced in the paper FarSLIP: Discovering…
62 results
MGRS-200k Dataset MGRS-200k is the first multi-granularity remote sensing (RS) image-text dataset, introduced in the paper FarSLIP: Discovering…
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with a Temporal-Aware Multimodal Model 📄 Paper (ICLR…
S1-MMAlign A Large-Scale Multi-Disciplinary Scientific Multimodal Dataset S1-MMAlign is a large-scale, multi-disciplinary multimodal dataset…
STRIDE-QA Dataset 📦 Dataset STRIDE-QA is a large-scale visual question answering (VQA) dataset for physically grounded spatiotemporal reasoning…
Cambrian-Alignment Dataset Please see paper & website for more information: Overview Cambrian-Alignment is an question-answering alignment dataset…
MAmmoTH-VL-Instruct-12M 🏠 Homepage 🤖 MAmmoTH-VL-8B 💻 Code 📄 Arxiv 📕 PDF 🖥️ Demo Introduction Our simple yet scalable…
NVSpeech Dataset Overview The NVSpeech dataset provides extensive annotations of paralinguistic vocalizations for Mandarin Chinese speech, aimed at…
OpenDialog OpenDialog is a 6.8k hours spoken dialogue dataset, introduced in the paper ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation…
ParCzech4Speech (Sentence-Segmented Variant) Dataset Summary ParCzech4Speech (Sentence-Segmented Variant) is a large-scale Czech speech dataset based…
PubTables-v2 PubTables-v2 is a new large-scale dataset for full-page and multi-page table extraction. Official dataset evaluation scripts and…
🎨 Danbooru2024 Webp 4MPixel Dataset 📊 Dataset Overview The Danbooru2024-Webp dataset is a comprehensive collection focused on animation…
Orient Anything V2 Dataset Project Page Paper GitHub Orient Anything V2 is an enhanced foundation model for unified…
FRoM-W1: Towards General Humanoid Whole-Body Control with Language Instructions The Humanoid Intelligence Team from FudanNLP and OpenMOSS…
Nano3D-Edit-100k This dataset is the official data release for Nano3D, a training-free framework for precise and coherent 3D…