Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
VisionArena-Bench: An automatic eval pipeline to estimate model preference rankings An automatic benchmark of 500 diverse user prompts that can be…
VisionArena-Bench: An automatic eval pipeline to estimate model preference rankings An automatic benchmark of 500 diverse user prompts that can be used to cheaply approximate Chatbot Arena model rankings via automatic benchmarking with VLM as a judge. Dataset Sources Repository: Paper: Automatic Evaluation Code: Coming Soon! Dataset Structure question id: The unique hash representing the… See the full description on the dataset page:
Source: Hugging Face Hub (lmarena-ai/vision-arena-bench-v0.1). Metadata imported from the dataset’s Hub tags.