Skip to content
Advertisement

Companies

554 results

SplxAI

Evaluation & Benchmarking

AI security company (Agentic Radar) providing automated red teaming and runtime protection across the AI lifecycle; acquired by Zscaler in…

Text

Repello AI

Evaluation & Benchmarking

AI application security company providing ARTEMIS, an automated red-teaming tool with attack patterns spanning text, image, and audio inputs.

Multimodal

WitnessAI

Evaluation & Benchmarking

Mountain View, CA-based AI governance and security platform providing visibility, policy control, and runtime protection for enterprise AI and agent…

Text

Noma Security

Evaluation & Benchmarking

AI security posture management company providing agentic access control, runtime protection, and red-teaming for enterprise AI systems.

Text

Straiker

Evaluation & Benchmarking

AI agent security company providing Discover, Ascend, and Defend products covering agent visibility, adversarial red-teaming, and runtime defense.

Text

Lasso Security

Evaluation & Benchmarking

AI agent security platform combining discovery, adversarial red teaming with thousands of attack techniques, and runtime protection.

Text

Promptfoo

Evaluation & Benchmarking

Open-source-rooted AI security testing platform for automated red teaming, guardrails, and LLM evaluation, used by over 300,000 developers; acquired by…

Text

Apollo Research

Evaluation & Benchmarking

London-based AI safety research organization (structured as a public benefit corporation) studying scheming and deceptive behavior in frontier AI models…

Text

METR

Evaluation & Benchmarking

Berkeley, CA-based nonprofit (Model Evaluation & Threat Research) conducting autonomous-capability evaluations of frontier AI models in partnership with major AI…

Text

Epoch AI

Evaluation & Benchmarking

AI research institute tracking model capability, compute, and infrastructure trends, including the Epoch Capabilities Index benchmark of frontier-model progress.

Text

Center for AI Safety

Evaluation & Benchmarking

San Francisco nonprofit (CAIS), led by Dan Hendrycks, that develops AI safety benchmarks including the MASK honesty benchmark and AgentHarm.

Text

Artificial Analysis

Evaluation & Benchmarking

London-based independent AI model benchmarking organization publishing the Artificial Analysis Intelligence Index and provider performance comparisons across cost, speed, and…

Multimodal

HackerOne

Evaluation & Benchmarking

San Francisco bug-bounty and security-research platform (founded 2012) whose AI Red Teaming service pairs vetted security researchers with AI-driven adversarial…

Text

Parea AI

Evaluation & Benchmarking

Y Combinator-backed AI testing and evaluation platform for experiment tracking, observability, and human review of LLM applications.

Text

Holistic AI

Evaluation & Benchmarking

AI governance platform offering an AI red-teaming module — dynamic adversarial testing, jailbreak-resistance checks, and prompt-injection detection — alongside 40+…

Text

Weights & Biases

Evaluation & Benchmarking

MLOps company (W&B) whose Weave product provides LLM agent evaluation — a flexible evaluation framework, leaderboards, and pre-built safety/quality scorers…

MultimodalText

Caruso

Data Marketplace

Connected-car data marketplace aggregating standardized telematics, diagnostic, and repair data from major automakers via a single API.

Sensor / Time-series

High Mobility

Proprietary Data Providers

Multi-brand connected-vehicle data API streaming raw location, diagnostic, EV-charging, and crash-event data from 22+ automakers.

Sensor / Time-series

Smartcar

Proprietary Data Providers

Vehicle data API unifying telematics, EV/energy, and diagnostic signals across 40+ automakers behind a single integration.

Sensor / Time-series

Geotab

Proprietary Data Providers

Commercial fleet telematics provider collecting GPS, engine-diagnostic, and EV-battery sensor data, with a Data Connector and third-party Marketplace for downstream…

Sensor / Time-series

CARFAX

Proprietary Data Providers

Vehicle history data provider aggregating over 35 billion ownership, accident, title, and maintenance records from 151,000+ sources.

Tabular

Arize AI

Evaluation & Benchmarking

Berkeley, CA-based AI observability and evaluation company behind Arize AX and the open-source Phoenix project.

Text

LMArena

Evaluation & Benchmarking

Community-driven LLM leaderboard originating from UC Berkeley's LMSYS project, now operated by the company Arena (arena.ai) with an enterprise evaluation…

ImageText

Confident AI

Evaluation & Benchmarking

Where AI Quality is Standardized. Not Improvised.

Text
Advertisement