Skip to main content
📖 The AI Tool Bible

Evaluation

Observability, prompt testing, and quality scoring.

59 tools

Why it matters

Evaluation is the discipline most underinvested in by AI product teams. Choosing an eval tool early is much cheaper than retrofitting one when an LLM regression hits production.

What's in here

Spans full eval + observability platforms (Braintrust, LangSmith), prompt management (Humanloop, PromptLayer), ML-broad tracking with LLM features (Weights & Biases), and proxy-based observability (Helicone).

How to pick

Pick Braintrust or LangSmith for full eval + observability. Pick Humanloop if PMs need to edit prompts. Pick Helicone for a one-line install on existing OpenAI/Claude code. Pick Patronus for automated hallucination/safety evals at scale.

LL

LLMEval

Evaluation · Multi-model
7.2

Open academic benchmark suite for stress-testing LLMs on contamination-resistant, domain-specific tasks.

Free· Free; open-source academic benchmarksllm-benchmarkingacademic-evaluation
PR

Promptfoo

Evaluation · Multi-model
7.2

Open-source eval and red-teaming framework for LLM apps, prompts, and RAG pipelines.

Freemium· Community: Free · Enterprise: Custom · On-Premise: Customllm-evalsred-teaming
WA

Weco AI

Evaluation · Multi-model (LLM + AIDE tree search)
7.2

Autoresearch engine that iteratively rewrites code to optimize against a numeric evaluation metric.

Freemium· Open-source CLI; hosted/commercial pricing not publishedcode-optimizationgpu-kernel-tuning
AL

AlpacaEval

Evaluation · GPT-4 Preview (Nov 2024) as annotator
7.1

Automatic LLM evaluator and leaderboard that benchmarks instruction-following with length-controlled win rates.

Free· Free and open-source; pay only for the underlying OpenAI annotator API callsllm-benchmarkinginstruction-following eval
AR

Arthur

Evaluation · Multi-model
7.1

Open-source toolkit for testing, tracing, and monitoring production AI agents.

Freemium· Free: $0/mo · Premium: $60/mo · Enterprise: Customagent-evaluationprompt-management
FA

Fiddler AI

Evaluation · Fiddler Centor (proprietary evaluators)
7.1

Enterprise AI observability and guardrails platform for monitoring agents, LLMs, and ML models in production.

Enterprise· Free: Free · Developer: $0.002 per trace · Enterprise: Contact salesllm-observabilityagent-monitoring
MA

Maxim AI

Evaluation · Multi-model
7.1

End-to-end evaluation, simulation, and observability platform for shipping production-grade AI agents.

Freemium· Developer: Free · Professional: $29 /seat /month · Business: $49 /seat /month · Enterprise: Customagent-evaluationllm-observability
PF

Prompt Foundry

Evaluation · OpenAI + Anthropic (multi-model)
7.1

Prompt management and side-by-side LLM evaluation for OpenAI and Anthropic models.

Freemium· Free tier (10 prompts, 500 evals/mo); Pro $15/user/mo; Enterprise customprompt-managementmodel-comparison
RF

Respan (formerly Keywords AI)

Evaluation · Multi-model (500+ via gateway)
7.1

LLM engineering platform combining a multi-model gateway with tracing, evals, and prompt management.

Freemium· Free tier; paid plans (pricing not public); enterprise on requestllm-observabilityprompt-management
SL

SEAL Leaderboard

Evaluation · Multi-model (GPT, Claude, Gemini, Llama, etc.)
7.1

Private, expert-graded leaderboards from Scale AI that rank frontier LLMs on domains contaminated public benchmarks can no longer measure.

Free· Free to view; paid custom evals via Scale enterprise salesmodel-selectionbenchmark-tracking
VI

VisualWebArena

Evaluation · Model-agnostic (GPT-4V, Gemini, Claude, open VLMs)
7.1

Open benchmark for evaluating multimodal web agents on realistic visual browsing tasks.

Free· Free and open source (MIT-style research release)multimodal-agent-evalweb-browsing-benchmark
LA

LangFast

Evaluation · Multi-model
7.0

No-signup LLM playground for testing, comparing, and versioning prompts against your own API keys.

Paid· One-time lifetime ~$60-$120; 14-day money-backprompt-testingprompt-versioning
LL

llmfit

Evaluation · Multi-model
7.0

Terminal tool that scores hundreds of open LLMs against your actual CPU, RAM, and GPU and tells you which ones will run well.

Free· Free, MIT-licensedlocal-llm-selectionhardware-benchmarking
OL

OlympicArena

Evaluation · GPT-4o, Claude-3.5-Sonnet, Doubao-Pro-32k, DeepSeek-Coder-V2, Qwen2-72B-Instruct
7.0

Olympiad-level multi-discipline benchmark for stress-testing reasoning in LLMs and multimodal models.

Free· Free, open-source research benchmarkllm-evaluationmultimodal-eval
PH

Phoenix

Evaluation · Multi-model
7.0

Open-source LLM and agent observability platform with tracing, evals, and experimentation built on OpenTelemetry.

Freemium· AX Free: Free · AX Pro: $50 · AX Enterprise: Customllm-tracingagent-debugging
SU

Superwise

Evaluation · Multi-model
7.0

Agentic management platform for runtime guardrails, policy enforcement, and observability across LLM agents.

Freemium· Starter: Free · Pro+: ? · Enterprise: ?llm-guardrailsai-governance
AG

Agenta

Evaluation · Multi-model
6.9

Open-source LLMOps platform for prompt engineering, evaluation, and observability in one workspace.

Freemium· Hobby: $0 forever · Pro: $29 /month · Business: $299 /month · Enterprise: Custom · Open source: Free foreverprompt-engineeringllm-evaluation
CO

CompassRank

Evaluation · Multi-model
6.9

Public leaderboard from the OpenCompass project ranking open and closed LLMs across 100+ benchmarks.

Free· Free leaderboard; OpenCompass toolkit is Apache 2.0 open sourcellm-benchmarkingmodel-selection
IN

InfiBench

Evaluation
6.9

Stack Overflow-derived benchmark for evaluating code LLMs on real-world programming questions.

Free· Free and open source (CC BY-SA 4.0)code-llm-evalmodel-benchmarking
MI

MixEval

Evaluation · GPT-3.5-Turbo-0125, GPT-4o-2024-05-13, Claude 3.5 Sonnet, MixEval, MixEval-Hard
6.9

Dynamic LLM benchmark that mixes web queries with existing datasets to mirror Chatbot Arena rankings at a fraction of the cost.

Free· Free and open sourcellm-benchmarkingmodel-ranking
AA

Arena AI

Evaluation · Multi-model
6.8

Head-to-head LLM battle arena with a public leaderboard for ranking AI models.

Free· Free to use; no public paid tier listedllm-benchmarkingmodel-comparison
AA

Artificial Analysis

Evaluation · Multi-model
6.8

Independent benchmarking platform comparing AI models and inference providers across intelligence, speed, and cost.

Freemium· Pro: $417/month per seat · Enterprise: Custom pricingmodel-benchmarkingprovider-comparison
CT

Cleanlab TLM

Evaluation · Multi-model (wraps any LLM)
6.8

Trustworthiness scoring layer that flags LLM hallucinations in real time.

Freemium· Free tier for evaluation; usage-based API pricing; enterprise/private deployment via saleshallucination-detectionrag-evaluation
PA

Parea AI

Evaluation · Multi-model
6.8

LLM evaluation, observability, and prompt management platform for teams shipping production AI apps.

Freemium· Free (2 seats, 3k logs/mo); Team $150/mo; Enterprise customllm-evaluationprompt-management