Skip to content
Advertisement
ImageMultimodalText

ULVR-filtered

ULVR-filtered Filtered subset of RuoliuYang/ULVR v2 clean: the 101,951 training samples that Qwen2.5-VL-7B-Instruct answered incorrectly given only…

ULVR-filtered Filtered subset of RuoliuYang/ULVR v2 clean: the 101,951 training samples that Qwen2.5-VL-7B-Instruct answered incorrectly given only input image, but correctly once the intermediate image were also provided (judged by Qwen3-VL-32B-Instruct). Same schema / subsets / train-split structure as the source. subset rows scene graph 3522 edge 1394 depth 537 segmentation 1328 bbox highlight 15186 bbox crop 15260 text cot 27158 helper interleaved… See the full description on the dataset page:

Source: Hugging Face Hub (williamium/ULVR-filtered). Metadata imported from the dataset’s Hub tags.

Advertisement