Skip to content
Advertisement

General

2,594 results

Lasso Security

Evaluation & Benchmarking

AI agent security platform combining discovery, adversarial red teaming with thousands of attack techniques, and runtime protection.

Text

Promptfoo

Evaluation & Benchmarking

Open-source-rooted AI security testing platform for automated red teaming, guardrails, and LLM evaluation, used by over 300,000 developers; acquired by…

Text

Apollo Research

Evaluation & Benchmarking

London-based AI safety research organization (structured as a public benefit corporation) studying scheming and deceptive behavior in frontier AI models…

Text

METR

Evaluation & Benchmarking

Berkeley, CA-based nonprofit (Model Evaluation & Threat Research) conducting autonomous-capability evaluations of frontier AI models in partnership with major AI…

Text

Epoch AI

Evaluation & Benchmarking

AI research institute tracking model capability, compute, and infrastructure trends, including the Epoch Capabilities Index benchmark of frontier-model progress.

Text

Center for AI Safety

Evaluation & Benchmarking

San Francisco nonprofit (CAIS), led by Dan Hendrycks, that develops AI safety benchmarks including the MASK honesty benchmark and AgentHarm.

Text

Artificial Analysis

Evaluation & Benchmarking

London-based independent AI model benchmarking organization publishing the Artificial Analysis Intelligence Index and provider performance comparisons across cost, speed, and…

Multimodal

HackerOne

Evaluation & Benchmarking

San Francisco bug-bounty and security-research platform (founded 2012) whose AI Red Teaming service pairs vetted security researchers with AI-driven adversarial…

Text

Parea AI

Evaluation & Benchmarking

Y Combinator-backed AI testing and evaluation platform for experiment tracking, observability, and human review of LLM applications.

Text

Holistic AI

Evaluation & Benchmarking

AI governance platform offering an AI red-teaming module — dynamic adversarial testing, jailbreak-resistance checks, and prompt-injection detection — alongside 40+…

Text

Weights & Biases

Evaluation & Benchmarking

MLOps company (W&B) whose Weave product provides LLM agent evaluation — a flexible evaluation framework, leaderboards, and pre-built safety/quality scorers…

MultimodalText

Arize AI

Evaluation & Benchmarking

Berkeley, CA-based AI observability and evaluation company behind Arize AX and the open-source Phoenix project.

Text

LMArena

Evaluation & Benchmarking

Community-driven LLM leaderboard originating from UC Berkeley's LMSYS project, now operated by the company Arena (arena.ai) with an enterprise evaluation…

ImageText

Confident AI

Evaluation & Benchmarking

Where AI Quality is Standardized. Not Improvised.

Text

Ragas

Evaluation & Benchmarking

Open-source RAG-evaluation framework, maintained under the ExplodingGradients/VibrantLabs project, providing automated metrics and synthetic evaluation-dataset generation.

Text

Vals AI

Evaluation & Benchmarking

Independent Evaluation, Unbiased Benchmarks

Text

Athina AI

Evaluation & Benchmarking

Ship AI to prod 10x faster

Text

HoneyHive

Evaluation & Benchmarking

The observability layer for production agents

Text

Freeplay

Evaluation & Benchmarking

The ops platform for AI engineering teams

Text

Openlayer

Evaluation & Benchmarking

Y Combinator-backed AI evaluation and governance platform running 100+ automated tests for issues such as prompt injection, PII leakage, and…

Text

Maxim AI

Evaluation & Benchmarking

Evaluation and observability platform (operated by H3 Labs Inc.) covering prompt experimentation, agent simulation, and production monitoring.

Text

Comet

Evaluation & Benchmarking

New York-based MLOps company whose Opik product is an open-source LLM observability and evaluation platform with 30+ built-in metrics.

Text

LangChain

Evaluation & Benchmarking

San Francisco company behind the LangChain framework and LangSmith, an agent and LLM observability/evaluation platform.

Text

Giskard

Evaluation & Benchmarking

Paris-based AI red-teaming and testing company providing black-box vulnerability scanning for conversational AI agents.

Text
Advertisement