Weights & Biases
Evaluation & BenchmarkingMLOps company (W&B) whose Weave product provides LLM agent evaluation — a flexible evaluation framework, leaderboards, and pre-built safety/quality scorers…
MultimodalTextThis category covers organisations that measure how AI systems behave. It spans four related activities: evaluation frameworks and platforms that score model and agent outputs against datasets, rubrics, and model-based judges; observability products that trace production traffic and run continuous online evaluation; adversarial testing and red-teaming, both automated attack tooling and human researcher networks, aimed at jailbreaks, prompt injection, data leakage, and agent misuse; and independent benchmarking and safety-research organisations, including nonprofits, that publish comparative indices and capability or risk evaluations rather than sell a product.Approach the category with a clear question, because the vendors answer different ones. Teams shipping an application usually need offline evaluation wired into continuous integration plus production tracing; teams releasing a model need capability and safety evaluations they can defend externally; teams under regulatory pressure need evidence that is reproducible and auditable. Questions that separate vendors include whether evaluation datasets are customer-owned or vendor-supplied and how contamination is avoided; how model-based judges are themselves validated, and what happens when the judge model is updated; whether the platform handles multi-step agent traces or only single-turn completions; how findings are prioritised rather than merely counted; whether a third party could reproduce the result; and, for independent benchmark publishers, how they are funded and whether the organisations they score are also customers.Tooling that reports latency, cost, and error rates without scoring output quality is general application monitoring. Governance and compliance is the adjacent category, concerned with policy, documentation, and regulatory obligation rather than with measured behaviour.
43 results
MLOps company (W&B) whose Weave product provides LLM agent evaluation — a flexible evaluation framework, leaderboards, and pre-built safety/quality scorers…
MultimodalTextBoston-area AI security company publishing the LLM Safety Leaderboard and providing agent red-teaming, guardrails, and policy-enforcement products.
TextAI application security company providing ARTEMIS, an automated red-teaming tool with attack patterns spanning text, image, and audio inputs.
MultimodalAI security posture management company providing agentic access control, runtime protection, and red-teaming for enterprise AI systems.
TextAI agent security platform combining discovery, adversarial red teaming with thousands of attack techniques, and runtime protection.
TextLondon-based AI safety research organization (structured as a public benefit corporation) studying scheming and deceptive behavior in frontier AI models…
TextSan Francisco nonprofit (CAIS), led by Dan Hendrycks, that develops AI safety benchmarks including the MASK honesty benchmark and AgentHarm.
TextLondon-based independent AI model benchmarking organization publishing the Artificial Analysis Intelligence Index and provider performance comparisons across cost, speed, and…
MultimodalAI governance platform offering an AI red-teaming module — dynamic adversarial testing, jailbreak-resistance checks, and prompt-injection detection — alongside 40+…
TextNew York AI red-teaming company founded by Leonard Tang, Steve Li, Richard Zhu, and Alex Gu, building automated adversarial testing…
TextSeattle-founded AI security company (Guardian, Recon, Layer) providing automated red teaming and model-security scanning; acquired by Palo Alto Networks in…
TextAustin, TX-based AI security company covering AI asset discovery, supply-chain security, adversarial attack simulation, and runtime defense.
TextAI security research and red-teaming company providing continuous adversarial testing and runtime protection for AI coding agents and applications.
TextAI security company operating Arena, a large-scale adversarial red-teaming network, plus the Shade red-teaming tool and Cygnal runtime protection.
Text