Evaluation & Benchmarking
This category covers organisations that measure how AI systems behave. It spans four related activities: evaluation frameworks and platforms that score model and agent outputs against datasets, rubrics, and model-based judges; observability products that trace production traffic and run continuous online evaluation; adversarial testing and red-teaming, both automated attack tooling and human researcher networks, aimed at jailbreaks, prompt injection, data leakage, and agent misuse; and independent benchmarking and safety-research organisations, including nonprofits, that publish comparative indices and capability or risk evaluations rather than sell a product.Approach the category with a clear question, because the vendors answer different ones. Teams shipping an application usually need offline evaluation wired into continuous integration plus production tracing; teams releasing a model need capability and safety evaluations they can defend externally; teams under regulatory pressure need evidence that is reproducible and auditable. Questions that separate vendors include whether evaluation datasets are customer-owned or vendor-supplied and how contamination is avoided; how model-based judges are themselves validated, and what happens when the judge model is updated; whether the platform handles multi-step agent traces or only single-turn completions; how findings are prioritised rather than merely counted; whether a third party could reproduce the result; and, for independent benchmark publishers, how they are funded and whether the organisations they score are also customers.Tooling that reports latency, cost, and error rates without scoring output quality is general application monitoring. Governance and compliance is the adjacent category, concerned with policy, documentation, and regulatory obligation rather than with measured behaviour.
43 results
Fiddler AI
Evaluation & BenchmarkingAI observability and governance company (founded by Krishna Gade and Amit Paka) providing continuous evaluation, guardrails, and compliance monitoring for…
TabularTextGuardrails AI
Evaluation & BenchmarkingAI reliability platform providing an open-source validation framework plus Snowglobe, a synthetic-data and evaluation-dataset generation tool.
TextPatronus AI
Evaluation & BenchmarkingSan Francisco AI evaluation company behind the Lynx hallucination-detection model, the FinanceBench benchmark, and the GLIDER evaluation model.
TextBraintrust
Evaluation & BenchmarkingThe AI observability platform for building quality AI products
Text