Skip to content
Advertisement

CARV: A Diagnostic Benchmark for Compositional Analogical Reasoning in Multimodal LLMs Authors: Yongkang Du, Xiaohan Zou, Minhao Cheng, Lu Lin · Pennsylvania State University Dataset Description CARV evaluates whether multimodal LLMs can compose transformation rules from multiple image pairs via logical set operations. Given n context pairs each depicting an atomic visual change, the model must synthesize a new rule through Union (∪), Intersection (∩), or… See the full description on the dataset page:

Source: Hugging Face Hub (duyongka/CARV). Metadata imported from the dataset’s Hub tags.

Advertisement