Text
2,175 results
Lasso Security
Evaluation & BenchmarkingAI agent security platform combining discovery, adversarial red teaming with thousands of attack techniques, and runtime protection.
TextApollo Research
Evaluation & BenchmarkingLondon-based AI safety research organization (structured as a public benefit corporation) studying scheming and deceptive behavior in frontier AI models…
TextCenter for AI Safety
Evaluation & BenchmarkingSan Francisco nonprofit (CAIS), led by Dan Hendrycks, that develops AI safety benchmarks including the MASK honesty benchmark and AgentHarm.
TextHolistic AI
Evaluation & BenchmarkingAI governance platform offering an AI red-teaming module — dynamic adversarial testing, jailbreak-resistance checks, and prompt-injection detection — alongside 40+…
TextWeights & Biases
Evaluation & BenchmarkingMLOps company (W&B) whose Weave product provides LLM agent evaluation — a flexible evaluation framework, leaderboards, and pre-built safety/quality scorers…
MultimodalText