Phoenix Evals

arize-ai/phoenix/.agents/skills/phoenix-evals

by arize-ai4b1fa58bc7d6Apache-2.011K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated today

Build and run evaluators for AI/LLM applications using Phoenix.

Instructions onlyAI & Agents
AI-generated overview

Guides building and running evaluators for AI/LLM applications with Phoenix, covering code and LLM judges, datasets, experiments, and CI gating.

What it does
This skill provides reference documentation for building evaluators for AI and LLM applications using Phoenix. It covers setup in Python and TypeScript, choosing judge models, pre-built and custom code or LLM evaluators, RAG evaluation, datasets and synthetic data, running experiments, validating evaluator accuracy against human labels, tracing and sampling, error analysis, and production guardrails. It is instructions only and produces no scripts or files itself.
When to use it
Use it when you need to create, run, or validate evaluators for an LLM application, set up Phoenix tracing and error analysis, build evaluation datasets or experiments, or gate CI on evaluation results.
Requirements
Requires a Phoenix server. Python work needs the phoenix and openai packages; TypeScript work needs @arizeai/phoenix-client. Ships no scripts; it is reference documentation only.

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

TaskFiles
Setupsetup-python [blocked], setup-typescript [blocked]
Decide what to evaluateevaluators-overview [blocked]
Choose a judge modelfundamentals-model-selection [blocked]
Use pre-built evaluatorsevaluators-pre-built [blocked]
Build code evaluatorevaluators-code-python [blocked], evaluators-code-typescript [blocked]
Build LLM evaluatorevaluators-llm-python [blocked], evaluators-llm-typescript [blocked], evaluators-custom-templates [blocked]
Batch evaluate DataFrameevaluate-dataframe-python [blocked]
Run experimentexperiments-running-python [blocked], experiments-running-typescript [blocked]
Run evals in a test runner (CI gate)integrations-pytest [blocked], integrations-vitest-jest [blocked]
Create datasetexperiments-datasets-python [blocked], experiments-datasets-typescript [blocked]
Generate synthetic dataexperiments-synthetic-python [blocked], experiments-synthetic-typescript [blocked]
Validate evaluator accuracyvalidation [blocked], validation-evaluators-python [blocked], validation-evaluators-typescript [blocked]
Export spansobserve-tracing-setup [blocked]
Write a span filter (SpanQuery().where)filter-expressions [blocked]
Sample traces for reviewobserve-sampling-python [blocked], observe-sampling-typescript [blocked]
Analyze errorserror-analysis [blocked], error-analysis-multi-turn [blocked], axial-coding [blocked]
RAG evalsevaluators-rag [blocked]
Avoid common mistakescommon-mistakes-python [blocked], fundamentals-anti-patterns [blocked]
Productionproduction-overview [blocked], production-guardrails [blocked], production-continuous [blocked]

Workflows

Starting Fresh: observe-tracing-setup [blocked] → error-analysis [blocked] → axial-coding [blocked] → evaluators-overview [blocked]

Building Evaluator: fundamentals [blocked] → common-mistakes-python [blocked] → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag [blocked] → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

Gating CI: evaluators-{code|llm}-{python|typescript} → integrations-{pytest|vitest-jest} → production-continuous [blocked]

Production: production-overview [blocked] → production-guardrails [blocked] → production-continuous [blocked]

Reference Categories

PrefixDescription
fundamentals-*Types, scores, anti-patterns
observe-*Tracing, sampling
error-analysis-*Finding failures
axial-coding-*Categorizing failures
evaluators-*Code, LLM, RAG evaluators
experiments-*Datasets, running experiments
integrations-*Run evals from test runners (pytest, Vitest, Jest) as a CI gate
validation-*Validating evaluator accuracy against human labels
production-*CI/CD, monitoring

Key Principles

PrincipleAction
Error analysis firstCan't automate what you haven't observed
Custom > genericBuild from your failures
Code firstDeterministic before LLM
Validate judges>80% TPR/TNR
Binary > LikertPass/fail, not 1-5
Invariants gate, signals trendassert/expect hard invariants (CI red); log LLM-judge quality signals and gate the aggregate (acceptance criteria), not every case

Source and attribution

Source:arize-ai/phoenixin.agents/skills/phoenix-evalsat commit4b1fa58

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal