Eval Harness

affaan-m/ECC/skills/eval-harness

作者 affaan-mef648e01899ba3e8dc6371642deaaf64b4477775無授權條款275K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫4 天前更新

Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k and pass^k reliability. Use when defining pass/fail criteria for agent tasks, measuring agent reliability, building regression suites for prompt or agent changes, or benchmarking across model versions.

AI 產生的概覽

為 AI 編碼工作階段定義評估驅動開發:能力與回歸評估、評分器以及 pass@k 可靠性指標。

功能
提供在實作前定義能力評估與回歸評估的框架,包含評估定義、執行報告與儲存配置的範本。它說明四種評分器類型(程式碼式、模型式、規則式與人工式)以及 pass@k 和 pass^k 等指標。它也說明本機工具的用途,包括膠囊分組,以及因缺少作業系統隔離而拒絕執行候選方案。
適用情境
適用於為代理任務設定通過/失敗標準、衡量代理可靠性、為提示詞或代理變更建立回歸測試套件,或跨模型版本進行基準測試。
執行需求
僅為說明性內容,技能不附帶指令碼。它引用 Bash、Read、Write、Edit、Grep 與 Glob 工具,並舉例提到 Node.js 工具路徑與 npm test 指令。

Eval Harness Skill

A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.

When to Activate

  • Setting up eval-driven development (EDD) for AI-assisted workflows
  • Defining pass/fail criteria for Claude Code task completion
  • Measuring agent reliability with pass@k metrics
  • Creating regression test suites for prompt or agent changes
  • Benchmarking agent performance across model versions

Philosophy

Eval-Driven Development treats evals as the "unit tests of AI development":

  • Define expected behavior BEFORE implementation
  • Run evals continuously during development
  • Track regressions with each change
  • Use pass@k metrics for reliability measurement

Eval Types

Capability Evals

Test if Claude can do something it couldn't before:

markdown
[CAPABILITY EVAL: feature-name]Task: Description of what Claude should accomplishSuccess Criteria:  - [ ] Criterion 1  - [ ] Criterion 2  - [ ] Criterion 3Expected Output: Description of expected result

Regression Evals

Ensure changes don't break existing functionality:

markdown
[REGRESSION EVAL: feature-name]Baseline: SHA or checkpoint nameTests:  - existing-test-1: PASS/FAIL  - existing-test-2: PASS/FAIL  - existing-test-3: PASS/FAILResult: X/Y passed (previously Y/Y)

Grader Types

1. Code-Based Grader

Deterministic checks using code:

bash
# Check if file contains expected patterngrep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"
# Check if tests passnpm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
# Check if build succeedsnpm run build && echo "PASS" || echo "FAIL"

2. Model-Based Grader

Use Claude to evaluate open-ended outputs:

markdown
[MODEL GRADER PROMPT]Evaluate the following code change:1. Does it solve the stated problem?2. Is it well-structured?3. Are edge cases handled?4. Is error handling appropriate?
Score: 1-5 (1=poor, 5=excellent)Reasoning: [explanation]

3. Human Grader

Flag for manual review:

markdown
[HUMAN REVIEW REQUIRED]Change: Description of what changedReason: Why human review is neededRisk Level: LOW/MEDIUM/HIGH

Metrics

pass@k

"At least one success in k attempts"

  • pass@1: First attempt success rate
  • pass@3: Success within 3 attempts
  • Typical target: pass@3 > 90%

pass^k

"All k trials succeed"

  • Higher bar for reliability
  • pass^3: 3 consecutive successes
  • Use for critical paths

Eval Workflow

1. Define (Before Coding)

markdown
## EVAL DEFINITION: feature-xyz
### Capability Evals1. Can create new user account2. Can validate email format3. Can hash password securely
### Regression Evals1. Existing login still works2. Session management unchanged3. Logout flow intact
### Success Metrics- pass@3 > 90% for capability evals- pass^3 = 100% for regression evals

2. Implement

Write code to pass the defined evals.

3. Evaluate

bash
# Run capability evals[Run each capability eval, record PASS/FAIL]
# Run regression evalsnpm test -- --testPathPattern="existing"
# Generate report

4. Report

markdown
EVAL REPORT: feature-xyz========================
Capability Evals:  create-user:     PASS (pass@1)  validate-email:  PASS (pass@2)  hash-password:   PASS (pass@1)  Overall:         3/3 passed
Regression Evals:  login-flow:      PASS  session-mgmt:    PASS  logout-flow:     PASS  Overall:         3/3 passed
Metrics:  pass@1: 67% (2/3)  pass@3: 100% (3/3)
Status: READY FOR REVIEW

Integration Patterns

Pre-Implementation

/eval define feature-name

Creates eval definition file at .claude/evals/feature-name.md

During Implementation

/eval check feature-name

Runs current evals and reports status

Post-Implementation

/eval report feature-name

Generates full eval report

Eval Storage

Store evals in project:

.claude/  evals/    feature-xyz.md      # Eval definition    feature-xyz.log     # Eval run history    baseline.json       # Regression baselines

Best Practices

  1. Define evals BEFORE coding - Forces clear thinking about success criteria
  2. Run evals frequently - Catch regressions early
  3. Track pass@k over time - Monitor reliability trends
  4. Use code graders when possible - Deterministic > probabilistic
  5. Human review for security - Never fully automate security checks
  6. Keep evals fast - Slow evals don't get run
  7. Version evals with code - Evals are first-class artifacts

Example: Adding Authentication

markdown
## EVAL: add-authentication
### Phase 1: Define (10 min)Capability Evals:- [ ] User can register with email/password- [ ] User can login with valid credentials- [ ] Invalid credentials rejected with proper error- [ ] Sessions persist across page reloads- [ ] Logout clears session
Regression Evals:- [ ] Public routes still accessible- [ ] API responses unchanged- [ ] Database schema compatible
### Phase 2: Implement (varies)[Write code]
### Phase 3: EvaluateRun: /eval check add-authentication
### Phase 4: ReportEVAL REPORT: add-authentication==============================Capability: 5/5 passed (pass@3: 100%)Regression: 3/3 passed (pass^3: 100%)Status: SHIP IT

Local Framework Utilities

The mechanical utilities ship in scripts/lib/eval-harness/:

sh
node scripts/eval-harness.js example
  • Capsule: hash-linked journal with five lineages and local integrity checks.
  • Inspection: source digests, validated variant paths, and syntactic warnings.
  • Replay: declared tools and content-addressed fixtures. Missing fixtures fail closed; SE3 and above are refused in replay. Record mode invokes the registered implementation, so only register trusted functions.
  • Receipt: offline verification of capsule and artifact bytes, with named checks.
  • Retrospective preparation: node scripts/eval-harness.js capsule group <dir> [<dir> ...] groups 1 to 100 explicitly selected, verified local capsule snapshots from one task family by declared harness version. Repeated snapshots count once; conflicting identities or invalid capsules reject the whole report. This is read-only record counting, with no new rollouts, scores or promotion. Use small, quiescent capsules. Payloads, directory arguments and raw run/capsule IDs are omitted, but task-family/version labels are verbatim and digest references are linkable; review them before sharing. Operational validation remains pending.

Candidate execution is disabled on every OS because no verified OS containment backend is implemented. gate run, runGate, runVariant, direct child launch, and the retired effect preload refuse with gate.isolation_required. No trust flag or caller-supplied executor can bypass the refusal. The example records that refusal and inspects source without executing or scoring it.

Do not present static warnings, a capsule receipt, or successful utility tests as candidate containment or promotion evidence. A future gate requires an independently reviewed OS boundary, protected checker and audit channels, and fatal baseline rejection. See docs/architecture/eval-harness-frameworks.md.

Product Evals (v1.8)

Use product evals when behavior quality cannot be captured by unit tests alone.

Grader Types

  1. Code grader (deterministic assertions)
  2. Rule grader (regex/schema constraints)
  3. Model grader (LLM-as-judge rubric)
  4. Human grader (manual adjudication for ambiguous outputs)

pass@k Guidance

  • pass@1: direct reliability
  • pass@3: practical reliability under controlled retries
  • pass^3: stability test (all 3 runs must pass)

Recommended thresholds:

  • Capability evals: pass@3 >= 0.90
  • Regression evals: pass^3 = 1.00 for release-critical paths

Eval Anti-Patterns

  • Overfitting prompts to known eval examples
  • Measuring only happy-path outputs
  • Ignoring cost and latency drift while chasing pass rates
  • Allowing flaky graders in release gates

Minimal Eval Artifact Layout

  • .claude/evals/<feature>.md definition
  • .claude/evals/<feature>.log run history
  • docs/releases/<version>/eval-summary.md release snapshot

來源與署名

來源:affaan-m/ECC位於skills/eval-harness提交ef648e0

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架