Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Code Leaderboard Results Paper RewardBench 2 Evaluation Dataset Card The RewardBench 2 evaluation dataset is the new version of RewardBench that is…
Code Leaderboard Results Paper RewardBench 2 Evaluation Dataset Card The RewardBench 2 evaluation dataset is the new version of RewardBench that is based on unseen human data and designed to be substantially more difficult! RewardBench 2 evaluates capabilities of reward models over the following categories: Factuality (NEW!): Tests the ability of RMs to detect hallucinations and other basic errors in completions. Precise Instruction Following (NEW!): Tests the ability of RMs… See the full description on the dataset page:
Source: Hugging Face Hub (allenai/reward-bench-2). Metadata imported from the dataset’s Hub tags.