Eval & Testing

LLM and AI-agent evaluation, prompt testing and benchmarking.

17 vendors in this category

Agenta

Agenta

Eval & Testing

Open-source LLM evaluation prompt management and observability platform

Horizontal (Industry Agnostic)

Germany

Arize AI

Arize AI

Eval & Testing

Agent observability evaluation and improvement platform with open-source Phoenix

Technology Horizontal (Industry Agnostic)

USA

Arthur AI

Arthur AI

Eval & Testing

ML monitoring fairness and LLM evaluation platform for enterprises

Horizontal (Industry Agnostic)

USA

Braintrust

Braintrust

Platforms & Products

AI observability platform for building quality AI products with evaluations, monitoring, and optimization tools

Horizontal (Industry Agnostic)
Confident AI (DeepEval)

Confident AI (DeepEval)

Eval & Testing

Open-source LLM evaluation framework with hosted observability platform

Financial Services Healthcare & Life Sciences

USA

Deepchecks

Deepchecks

Eval & Testing

Continuous validation and testing platform for ML models and LLM apps

Israel

Galileo

Galileo

Analytics & Conversation Intelligence

AI observability and evaluation platform that turns offline evals into production guardrails for AI systems

Horizontal (Industry Agnostic)
Giskard

Giskard

Eval & Testing

Open-source LLM testing and evaluation framework for quality and safety

Financial Services Healthcare & Life Sciences

France

LangWatch

LangWatch

Conversational & Voice QA

AI agent testing LLM evaluation and observability platform

Technology Horizontal (Industry Agnostic)

Netherlands

Maxim AI

Maxim AI

Eval & Testing

End-to-end evaluation and observability platform for AI agents

Technology Horizontal (Industry Agnostic)

India/USA

Openlayer

Openlayer

Eval & Testing

AI agent evaluation and stress-testing platform for pre-deployment

Retail & E-commerce Travel & Hospitality

USA

Opik

Opik

Open Source Projects

Open-source LLM evaluation platform for debugging, evaluating, and monitoring LLM applications and RAG systems

Horizontal (Industry Agnostic)
Patronus AI

Patronus AI

Eval & Testing

Automated LLM evaluation and security platform for regulated industries

Financial Services Technology

USA

Promptfoo

Promptfoo

Eval & Testing

Open-source AI security and testing platform for LLM vulnerabilities

Financial Services Healthcare & Life Sciences

USA

Ragas

Ragas

Eval & Testing

Open-source framework for evaluating RAG pipelines and LLM applications

Technology Horizontal (Industry Agnostic)

USA

T

TruLens / TruEra

Eval & Testing

Open-source LLM evaluation and tracing framework (TruEra acquired by Snowflake)

Horizontal (Industry Agnostic)

USA

Vellum

Vellum

Eval & Testing

LLM development platform with evaluation and testing capabilities

Horizontal (Industry Agnostic)

USA