Skip to content
Advertisement
ImageMultimodalText

MM-JudgeBench

MM-JudgeBench

MM-JudgeBench Dataset Summary MM-JudgeBench is a multilingual multimodal preference benchmark for evaluating vision-language judge and reward models. Each row contains an image reference, a query, two candidate responses, and a preference label. The dataset includes three configurations: m-vl-rewardbench m-opencqa m-mm-rewardbench Each configuration provides two splits: original reversed In the reversed split, the response order is swapped and the preference… See the full description on the dataset page:

Source: Hugging Face Hub (tahmedge/MM-JudgeBench). Metadata imported from the dataset’s Hub tags.

Advertisement