Skip to main content
๐Ÿ“– The AI Tool Bible

Arize AI alternatives

12 evaluation tools in the same lane as Arize AI, ranked by editorial score.

โ† Back to Arize AI
BR

Braintrust

Featured
Evaluation ยท Platform (any LLM)
8.9

Eval, monitor, and improve AI products end-to-end.

Freemiumยท Starter: $0 ยท Pro: $249 ยท Enterprise: Custom pricingevalsmonitoring
LA

LangSmith

Evaluation ยท Platform (any LLM)
8.7

LangChain's eval + observability platform.

Freemiumยท Developer: $0 ยท Plus: $39 ยท Enterprise: Custom pricingLLM tracingevals
WB

Weights & Biases

Evaluation ยท Platform (any LLM)
8.4

The ML experiment tracker, now with LLM eval features.

Freemiumยท Free: $0/mo ยท Pro: Starts at $60/month ยท Enterprise: Custom plans ยท Personal: $0/mo ยท Advanced Enterprise: Custom planML experimentsLLM eval
HE

Helicone

Evaluation ยท Platform (any LLM)
8.3

Open-source LLM observability โ€” one-line proxy install.

Freemiumยท Free 100k req/mo; Pro from $25/moobservabilitycost tracking
GI

Giskard

Evaluation ยท Multi-model
8.2

Continuous AI red teaming platform that stress-tests LLM agents for vulnerabilities before they hit production.

Freemiumยท Open-source free tier; Giskard Hub enterprise pricing on requestllm-red-teamingagent-security-testing
GE

Great Expectations

Evaluation
8.2

Open-source data quality framework for validating the datasets that feed your ML and analytics pipelines.

Freemiumยท Developer: Free ยท Team: Custom ยท Enterprise: Contact Salesdata-validationpipeline-testing
HU

Humanloop

Evaluation ยท Platform (any LLM)
8.2

Prompt management + evals for collaborative AI teams.

Paidยท From $200/mo teamprompt managementteam collab
LI

LiveBench

Evaluation ยท Multi-model
8.2

Contamination-free LLM benchmark that refreshes its questions monthly to keep frontier models honest.

Freeยท Free and open source; self-hosted evaluation runnerllm-benchmarkingmodel-selection
AA

Athina AI

Evaluation ยท Multi-model
8.1

Collaborative LLM evaluation and observability platform for teams shipping AI features to production.

Freemiumยท Starter free (10k logs/mo); Pro & Enterprise customllm-evaluationprompt-management
BF

Berkeley Function-Calling Leaderboard

Evaluation ยท Multi-model
8.1

Open benchmark from UC Berkeley that ranks LLMs on real-world tool-use and function-calling accuracy.

Freeยท Free and open source; you pay only for inference when reproducing runs.function-calling evaltool-use benchmarking
HO

HoneyHive

Evaluation ยท Multi-model
8.1

OpenTelemetry-native observability and evaluation platform for LLM agents in production.

Freemiumยท Free tier available; paid/enterprise tiers via salesagent-observabilityllm-evaluation
ML

MLflow

Evaluation ยท Multi-model
8.1

Open-source platform for tracking, evaluating, and deploying ML models and LLM applications.

Freeยท Free and open source (Apache 2.0); managed offering via Databricksllm-evaluationexperiment-tracking
Arize AI alternatives โ€” what to use instead in 2026 ยท The AI Tool Bible