Skip to content
Advertisement

Companies

554 results

Ragas

Evaluation & Benchmarking

Open-source RAG-evaluation framework, maintained under the ExplodingGradients/VibrantLabs project, providing automated metrics and synthetic evaluation-dataset generation.

Text

Vals AI

Evaluation & Benchmarking

Independent Evaluation, Unbiased Benchmarks

Text

Athina AI

Evaluation & Benchmarking

Ship AI to prod 10x faster

Text

HoneyHive

Evaluation & Benchmarking

The observability layer for production agents

Text

Freeplay

Evaluation & Benchmarking

The ops platform for AI engineering teams

Text

Openlayer

Evaluation & Benchmarking

Y Combinator-backed AI evaluation and governance platform running 100+ automated tests for issues such as prompt injection, PII leakage, and…

Text

Maxim AI

Evaluation & Benchmarking

Evaluation and observability platform (operated by H3 Labs Inc.) covering prompt experimentation, agent simulation, and production monitoring.

Text

Comet

Evaluation & Benchmarking

New York-based MLOps company whose Opik product is an open-source LLM observability and evaluation platform with 30+ built-in metrics.

Text

LangChain

Evaluation & Benchmarking

San Francisco company behind the LangChain framework and LangSmith, an agent and LLM observability/evaluation platform.

Text

Giskard

Evaluation & Benchmarking

Paris-based AI red-teaming and testing company providing black-box vulnerability scanning for conversational AI agents.

Text

Fiddler AI

Evaluation & Benchmarking

AI observability and governance company (founded by Krishna Gade and Amit Paka) providing continuous evaluation, guardrails, and compliance monitoring for…

TabularText

Vectara

Evaluation & Benchmarking

Palo Alto RAG platform that publishes the Hallucination Evaluation Model (HHEM) and an associated hallucination-leaderboard for LLMs.

Text

Guardrails AI

Evaluation & Benchmarking

AI reliability platform providing an open-source validation framework plus Snowglobe, a synthetic-data and evaluation-dataset generation tool.

Text

Lakera

Evaluation & Benchmarking

Zurich-founded generative-AI security company providing AI red teaming alongside runtime prompt-attack and data-leakage protection; acquired by Check Point Software in…

Text

Gray Swan AI

Evaluation & Benchmarking

AI security company operating Arena, a large-scale adversarial red-teaming network, plus the Shade red-teaming tool and Cygnal runtime protection.

Text

Haize Labs

Evaluation & Benchmarking

New York AI red-teaming company founded by Leonard Tang, Steve Li, Richard Zhu, and Alex Gu, building automated adversarial testing…

Text

Adversa AI

Evaluation & Benchmarking

AI security research and red-teaming company providing continuous adversarial testing and runtime protection for AI coding agents and applications.

Text

HiddenLayer

Evaluation & Benchmarking

Austin, TX-based AI security company covering AI asset discovery, supply-chain security, adversarial attack simulation, and runtime defense.

Text

Protect AI

Evaluation & Benchmarking

Seattle-founded AI security company (Guardian, Recon, Layer) providing automated red teaming and model-security scanning; acquired by Palo Alto Networks in…

Text

Patronus AI

Evaluation & Benchmarking

San Francisco AI evaluation company behind the Lynx hallucination-detection model, the FinanceBench benchmark, and the GLIDER evaluation model.

Text

Galileo

Evaluation & Benchmarking

Don't just monitor AI failures. Stop them.

Text

Braintrust

Evaluation & Benchmarking

The AI observability platform for building quality AI products

Text

TranscribeMe

Labeling & Annotation

Transcription company offering human-verified speech-to-text data and custom AI dataset creation for model training.

Audio

Datamundi

Labeling & Annotation

We Create High Quality Human Data To Fuel Your AI

AudioText
Advertisement