Multimodal
2,368 results
Recap-DataComp-1B
Dataset Card for Recap-DataComp-1B Recap-DataComp-1B is a large-scale image-text dataset that has been recaptioned using an advanced…
megalith-mdqa
Images from Megalith, synthetically captioned using Moondream, with the questions then transformed to short-form QA using an LLM.
FLUX-Reason-6M
FLUX-Reason-6M FLUX-Reason-6M is a massive, 6-million-scale text-to-image dataset engineered to instill complex reasoning capabilities in generative…
GSA_volc
GSA volc - GSA Embodied Perception Training Dataset Large-scale Grounding-Spatial-Affordance (GSA) training data for embodied perception Teacher…
MMStar
MMStar (Are We on the Right Way for Evaluating Large Vision-Language Models?) 🌐 Homepage 🤗 Dataset 🤗 Paper…
WebLINX-full
WebLINX: Real-World Website Navigation with Multi-Turn Dialogue WARNING: This is not the main WebLINX data card! You might…
chitralekha
Chitralekha Dataset Details Dataset Version Some of the fonts do not have proper letters/rendering of different telugu letter…
scallop_lr_teacher_640
FishingROV: scallop lr teacher 640 This dataset is an optimized derivative format generated for the FishingROV edge inference…
idl-wds
Dataset Card for Industry Documents Library (IDL) Dataset Summary Industry Documents Library (IDL) is a document dataset filtered…
QIN-LungCT-Seg (multi-site lung nodule segmentations)
QIN-LungCT-Seg (multi-site lung nodule segmentations)
synthetic-living-room-dataset-for-robotic-perception
Synthetic Living Room Dataset for Robotic Perception Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from…
LayeredFlow-Syn Extracted Ground Truth
LayeredFlow-Syn Extracted Ground Truth
AmbientCG Textures and HDRIs
AmbientCG Textures and HDRIs