Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
MM-JudgeBench Dataset Summary MM-JudgeBench is a multilingual multimodal preference benchmark for evaluating vision-language judge and reward models. Each row contains an image reference, a query, two candidate responses, and a preference label. The dataset includes three configurations: m-vl-rewardbench m-opencqa m-mm-rewardbench Each configuration provides two splits: original reversed In the reversed split, the response order is swapped and the preference… See the full description on the dataset page:
Source: Hugging Face Hub (tahmedge/MM-JudgeBench). Metadata imported from the dataset’s Hub tags.