Open X-Embodiment
1M+ real-robot trajectories from 20+ institutions.
Liars' Bench Liars' Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here. It was generated by using...…
Liars’ Bench Liars’ Bench is a benchmark for evaluating lie-detectors for language models. For details, see paper here. It was generated by using… four open-weight models: mistral-small-3.1-24b-instruct, llama-v3.3-70b-instruct, qwen-2.5-72b-instruct, gemma-3-27b-it across seven datasets including a control dataset (Alpaca). Our settings capture qualitatively different types of lies and vary along two dimensions: the model’s reason for lying the object of belief targeted by… See the full description on the dataset page:
Source: Hugging Face Hub (Cadenza-Labs/liars-bench). Metadata imported from the dataset’s Hub tags.