Phoenix Evals

arize-ai/phoenix/.agents/skills/phoenix-evals

作者 arize-ai4b1fa58bc7d6Apache-2.011K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫今天更新

Build and run evaluators for AI/LLM applications using Phoenix.

僅含說明AI & Agents
AI 產生的概覽

指導使用 Phoenix 為 AI/LLM 應用程式建置與執行評估器,涵蓋程式碼與 LLM 評判、資料集、實驗和 CI 閘門。

功能
此技能提供使用 Phoenix 為 AI 與 LLM 應用程式建置評估器的參考文件。內容涵蓋 Python 與 TypeScript 環境設定、評判模型選擇、預建及自訂的程式碼或 LLM 評估器、RAG 評估、資料集與合成資料、執行實驗、依人工標註驗證評估器準確度、追蹤與取樣、錯誤分析以及正式環境防護措施。它僅為說明文件,本身不產生指令碼或檔案。
適用情境
當你需要為 LLM 應用程式建立、執行或驗證評估器,設定 Phoenix 追蹤與錯誤分析,建置評估資料集或實驗,或依評估結果設定 CI 閘門時使用。
執行需求
需要 Phoenix 伺服器。Python 相關工作需要 phoenix 和 openai 套件;TypeScript 相關工作需要 @arizeai/phoenix-client。不附帶指令碼,僅為參考文件。

Phoenix Evals

Build evaluators for AI/LLM applications. Code first, LLM for nuance, validate against humans.

Quick Reference

TaskFiles
Setupsetup-python [blocked], setup-typescript [blocked]
Decide what to evaluateevaluators-overview [blocked]
Choose a judge modelfundamentals-model-selection [blocked]
Use pre-built evaluatorsevaluators-pre-built [blocked]
Build code evaluatorevaluators-code-python [blocked], evaluators-code-typescript [blocked]
Build LLM evaluatorevaluators-llm-python [blocked], evaluators-llm-typescript [blocked], evaluators-custom-templates [blocked]
Batch evaluate DataFrameevaluate-dataframe-python [blocked]
Run experimentexperiments-running-python [blocked], experiments-running-typescript [blocked]
Run evals in a test runner (CI gate)integrations-pytest [blocked], integrations-vitest-jest [blocked]
Create datasetexperiments-datasets-python [blocked], experiments-datasets-typescript [blocked]
Generate synthetic dataexperiments-synthetic-python [blocked], experiments-synthetic-typescript [blocked]
Validate evaluator accuracyvalidation [blocked], validation-evaluators-python [blocked], validation-evaluators-typescript [blocked]
Export spansobserve-tracing-setup [blocked]
Write a span filter (SpanQuery().where)filter-expressions [blocked]
Sample traces for reviewobserve-sampling-python [blocked], observe-sampling-typescript [blocked]
Analyze errorserror-analysis [blocked], error-analysis-multi-turn [blocked], axial-coding [blocked]
RAG evalsevaluators-rag [blocked]
Avoid common mistakescommon-mistakes-python [blocked], fundamentals-anti-patterns [blocked]
Productionproduction-overview [blocked], production-guardrails [blocked], production-continuous [blocked]

Workflows

Starting Fresh: observe-tracing-setup [blocked] → error-analysis [blocked] → axial-coding [blocked] → evaluators-overview [blocked]

Building Evaluator: fundamentals [blocked] → common-mistakes-python [blocked] → evaluators-{code|llm}-{python|typescript} → validation-evaluators-{python|typescript}

RAG Systems: evaluators-rag [blocked] → evaluators-code-* (retrieval) → evaluators-llm-* (faithfulness)

Gating CI: evaluators-{code|llm}-{python|typescript} → integrations-{pytest|vitest-jest} → production-continuous [blocked]

Production: production-overview [blocked] → production-guardrails [blocked] → production-continuous [blocked]

Reference Categories

PrefixDescription
fundamentals-*Types, scores, anti-patterns
observe-*Tracing, sampling
error-analysis-*Finding failures
axial-coding-*Categorizing failures
evaluators-*Code, LLM, RAG evaluators
experiments-*Datasets, running experiments
integrations-*Run evals from test runners (pytest, Vitest, Jest) as a CI gate
validation-*Validating evaluator accuracy against human labels
production-*CI/CD, monitoring

Key Principles

PrincipleAction
Error analysis firstCan't automate what you haven't observed
Custom > genericBuild from your failures
Code firstDeterministic before LLM
Validate judges>80% TPR/TNR
Binary > LikertPass/fail, not 1-5
Invariants gate, signals trendassert/expect hard invariants (CI red); log LLM-judge quality signals and gate the aggregate (acceptance criteria), not every case

來源與署名

來源:arize-ai/phoenix位於.agents/skills/phoenix-evals提交4b1fa58

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架