Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
PandaBench PandaBench is a comprehensive benchmark for evaluating Large Language Model (LLM) safety, focusing on jailbreak attacks, defense mechanisms, and evaluation methodologies. The PandaGuard framework architecture illustrating the end-to-end pipeline for LLM safety evaluation. The system connects three key components: Attackers, Defenders, and Judges. Dataset Description This repository contains the benchmark results from extensive evaluations of various… See the full description on the dataset page:
Source: Hugging Face Hub (Beijing-AISI/panda-bench). Metadata imported from the dataset’s Hub tags.