Skip to content
Advertisement

Evaluation & Benchmarking

This category covers organisations that measure how AI systems behave. It spans four related activities: evaluation frameworks and platforms that score model and agent outputs against datasets, rubrics, and model-based judges; observability products that trace production traffic and run continuous online evaluation; adversarial testing and red-teaming, both automated attack tooling and human researcher networks, aimed at jailbreaks, prompt injection, data leakage, and agent misuse; and independent benchmarking and safety-research organisations, including nonprofits, that publish comparative indices and capability or risk evaluations rather than sell a product.Approach the category with a clear question, because the vendors answer different ones. Teams shipping an application usually need offline evaluation wired into continuous integration plus production tracing; teams releasing a model need capability and safety evaluations they can defend externally; teams under regulatory pressure need evidence that is reproducible and auditable. Questions that separate vendors include whether evaluation datasets are customer-owned or vendor-supplied and how contamination is avoided; how model-based judges are themselves validated, and what happens when the judge model is updated; whether the platform handles multi-step agent traces or only single-turn completions; how findings are prioritised rather than merely counted; whether a third party could reproduce the result; and, for independent benchmark publishers, how they are funded and whether the organisations they score are also customers.Tooling that reports latency, cost, and error rates without scoring output quality is general application monitoring. Governance and compliance is the adjacent category, concerned with policy, documentation, and regulatory obligation rather than with measured behaviour.

43 results

Weights & Biases

Evaluation & Benchmarking

MLOps company (W&B) whose Weave product provides LLM agent evaluation — a flexible evaluation framework, leaderboards, and pre-built safety/quality scorers…

MultimodalText

Mindgard

Evaluation & Benchmarking

AI security company spun out of a decade of AI security research at Lancaster University, providing automated AI red teaming…

Text

Enkrypt AI

Evaluation & Benchmarking

Boston-area AI security company publishing the LLM Safety Leaderboard and providing agent red-teaming, guardrails, and policy-enforcement products.

Text

SplxAI

Evaluation & Benchmarking

AI security company (Agentic Radar) providing automated red teaming and runtime protection across the AI lifecycle; acquired by Zscaler in…

Text

Repello AI

Evaluation & Benchmarking

AI application security company providing ARTEMIS, an automated red-teaming tool with attack patterns spanning text, image, and audio inputs.

Multimodal

WitnessAI

Evaluation & Benchmarking

Mountain View, CA-based AI governance and security platform providing visibility, policy control, and runtime protection for enterprise AI and agent…

Text

Noma Security

Evaluation & Benchmarking

AI security posture management company providing agentic access control, runtime protection, and red-teaming for enterprise AI systems.

Text

Straiker

Evaluation & Benchmarking

AI agent security company providing Discover, Ascend, and Defend products covering agent visibility, adversarial red-teaming, and runtime defense.

Text

Lasso Security

Evaluation & Benchmarking

AI agent security platform combining discovery, adversarial red teaming with thousands of attack techniques, and runtime protection.

Text

Promptfoo

Evaluation & Benchmarking

Open-source-rooted AI security testing platform for automated red teaming, guardrails, and LLM evaluation, used by over 300,000 developers; acquired by…

Text

Apollo Research

Evaluation & Benchmarking

London-based AI safety research organization (structured as a public benefit corporation) studying scheming and deceptive behavior in frontier AI models…

Text

METR

Evaluation & Benchmarking

Berkeley, CA-based nonprofit (Model Evaluation & Threat Research) conducting autonomous-capability evaluations of frontier AI models in partnership with major AI…

Text

Epoch AI

Evaluation & Benchmarking

AI research institute tracking model capability, compute, and infrastructure trends, including the Epoch Capabilities Index benchmark of frontier-model progress.

Text

Center for AI Safety

Evaluation & Benchmarking

San Francisco nonprofit (CAIS), led by Dan Hendrycks, that develops AI safety benchmarks including the MASK honesty benchmark and AgentHarm.

Text

Artificial Analysis

Evaluation & Benchmarking

London-based independent AI model benchmarking organization publishing the Artificial Analysis Intelligence Index and provider performance comparisons across cost, speed, and…

Multimodal

HackerOne

Evaluation & Benchmarking

San Francisco bug-bounty and security-research platform (founded 2012) whose AI Red Teaming service pairs vetted security researchers with AI-driven adversarial…

Text

Parea AI

Evaluation & Benchmarking

Y Combinator-backed AI testing and evaluation platform for experiment tracking, observability, and human review of LLM applications.

Text

Holistic AI

Evaluation & Benchmarking

AI governance platform offering an AI red-teaming module — dynamic adversarial testing, jailbreak-resistance checks, and prompt-injection detection — alongside 40+…

Text

Haize Labs

Evaluation & Benchmarking

New York AI red-teaming company founded by Leonard Tang, Steve Li, Richard Zhu, and Alex Gu, building automated adversarial testing…

Text

Protect AI

Evaluation & Benchmarking

Seattle-founded AI security company (Guardian, Recon, Layer) providing automated red teaming and model-security scanning; acquired by Palo Alto Networks in…

Text

HiddenLayer

Evaluation & Benchmarking

Austin, TX-based AI security company covering AI asset discovery, supply-chain security, adversarial attack simulation, and runtime defense.

Text

Adversa AI

Evaluation & Benchmarking

AI security research and red-teaming company providing continuous adversarial testing and runtime protection for AI coding agents and applications.

Text

Gray Swan AI

Evaluation & Benchmarking

AI security company operating Arena, a large-scale adversarial red-teaming network, plus the Shade red-teaming tool and Cygnal runtime protection.

Text

Arize AI

Evaluation & Benchmarking

Berkeley, CA-based AI observability and evaluation company behind Arize AX and the open-source Phoenix project.

Text
Advertisement