Skip to content
Advertisement
ImageMultimodalText

ULVR_v2_clean

ULVR v2 clean Universal Latent Visual Reasoning training data, cleaned. 8 categories (subsets); each has train + validation splits. Every sample:…

ULVR v2 clean Universal Latent Visual Reasoning training data, cleaned. 8 categories (subsets); each has train + validation splits. Every sample: input image + question – assistant produces + intermediate visual step(s) + boxed{answer}. subset train validation text cot 333,911 3,533 bbox highlight 229,237 2,558 bbox crop 229,237 2,558 depth 40,000 25 edge 40,000 14 segmentation 40,000 326 helper interleaved 340,210 3,544 scene graph 40… See the full description on the dataset page: v2 clean.

Source: Hugging Face Hub (RuoliuYang/ULVR_v2_clean). Metadata imported from the dataset’s Hub tags.

Advertisement