Eval Harness

affaan-m/ECC/skills/eval-harness

作者 affaan-mef648e01899ba3e8dc6371642deaaf64b4477775无许可证275K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库4天前更新

Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k and pass^k reliability. Use when defining pass/fail criteria for agent tasks, measuring agent reliability, building regression suites for prompt or agent changes, or benchmarking across model versions.

AI 生成的概览

为 AI 编码会话定义评估驱动开发:能力与回归评估、评分器以及 pass@k 可靠性指标。

功能
提供在实现前定义能力评估与回归评估的框架,包含评估定义、运行报告和存储布局的模板。它描述四种评分器类型(基于代码、基于模型、规则和人工)以及 pass@k 和 pass^k 等指标。它还说明本地工具的用途,包括胶囊分组,以及因缺少操作系统隔离而拒绝执行候选方案。
适用场景
适用于为智能体任务设定通过/失败标准、衡量智能体可靠性、为提示词或智能体变更构建回归测试套件,或跨模型版本进行基准测试。
运行要求
仅为说明性内容,技能不附带脚本。它引用 Bash、Read、Write、Edit、Grep 和 Glob 工具,并举例提到 Node.js 工具路径和 npm test 命令。

Eval Harness Skill

A formal evaluation framework for Claude Code sessions, implementing eval-driven development (EDD) principles.

When to Activate

  • Setting up eval-driven development (EDD) for AI-assisted workflows
  • Defining pass/fail criteria for Claude Code task completion
  • Measuring agent reliability with pass@k metrics
  • Creating regression test suites for prompt or agent changes
  • Benchmarking agent performance across model versions

Philosophy

Eval-Driven Development treats evals as the "unit tests of AI development":

  • Define expected behavior BEFORE implementation
  • Run evals continuously during development
  • Track regressions with each change
  • Use pass@k metrics for reliability measurement

Eval Types

Capability Evals

Test if Claude can do something it couldn't before:

markdown
[CAPABILITY EVAL: feature-name]Task: Description of what Claude should accomplishSuccess Criteria:  - [ ] Criterion 1  - [ ] Criterion 2  - [ ] Criterion 3Expected Output: Description of expected result

Regression Evals

Ensure changes don't break existing functionality:

markdown
[REGRESSION EVAL: feature-name]Baseline: SHA or checkpoint nameTests:  - existing-test-1: PASS/FAIL  - existing-test-2: PASS/FAIL  - existing-test-3: PASS/FAILResult: X/Y passed (previously Y/Y)

Grader Types

1. Code-Based Grader

Deterministic checks using code:

bash
# Check if file contains expected patterngrep -q "export function handleAuth" src/auth.ts && echo "PASS" || echo "FAIL"
# Check if tests passnpm test -- --testPathPattern="auth" && echo "PASS" || echo "FAIL"
# Check if build succeedsnpm run build && echo "PASS" || echo "FAIL"

2. Model-Based Grader

Use Claude to evaluate open-ended outputs:

markdown
[MODEL GRADER PROMPT]Evaluate the following code change:1. Does it solve the stated problem?2. Is it well-structured?3. Are edge cases handled?4. Is error handling appropriate?
Score: 1-5 (1=poor, 5=excellent)Reasoning: [explanation]

3. Human Grader

Flag for manual review:

markdown
[HUMAN REVIEW REQUIRED]Change: Description of what changedReason: Why human review is neededRisk Level: LOW/MEDIUM/HIGH

Metrics

pass@k

"At least one success in k attempts"

  • pass@1: First attempt success rate
  • pass@3: Success within 3 attempts
  • Typical target: pass@3 > 90%

pass^k

"All k trials succeed"

  • Higher bar for reliability
  • pass^3: 3 consecutive successes
  • Use for critical paths

Eval Workflow

1. Define (Before Coding)

markdown
## EVAL DEFINITION: feature-xyz
### Capability Evals1. Can create new user account2. Can validate email format3. Can hash password securely
### Regression Evals1. Existing login still works2. Session management unchanged3. Logout flow intact
### Success Metrics- pass@3 > 90% for capability evals- pass^3 = 100% for regression evals

2. Implement

Write code to pass the defined evals.

3. Evaluate

bash
# Run capability evals[Run each capability eval, record PASS/FAIL]
# Run regression evalsnpm test -- --testPathPattern="existing"
# Generate report

4. Report

markdown
EVAL REPORT: feature-xyz========================
Capability Evals:  create-user:     PASS (pass@1)  validate-email:  PASS (pass@2)  hash-password:   PASS (pass@1)  Overall:         3/3 passed
Regression Evals:  login-flow:      PASS  session-mgmt:    PASS  logout-flow:     PASS  Overall:         3/3 passed
Metrics:  pass@1: 67% (2/3)  pass@3: 100% (3/3)
Status: READY FOR REVIEW

Integration Patterns

Pre-Implementation

/eval define feature-name

Creates eval definition file at .claude/evals/feature-name.md

During Implementation

/eval check feature-name

Runs current evals and reports status

Post-Implementation

/eval report feature-name

Generates full eval report

Eval Storage

Store evals in project:

.claude/  evals/    feature-xyz.md      # Eval definition    feature-xyz.log     # Eval run history    baseline.json       # Regression baselines

Best Practices

  1. Define evals BEFORE coding - Forces clear thinking about success criteria
  2. Run evals frequently - Catch regressions early
  3. Track pass@k over time - Monitor reliability trends
  4. Use code graders when possible - Deterministic > probabilistic
  5. Human review for security - Never fully automate security checks
  6. Keep evals fast - Slow evals don't get run
  7. Version evals with code - Evals are first-class artifacts

Example: Adding Authentication

markdown
## EVAL: add-authentication
### Phase 1: Define (10 min)Capability Evals:- [ ] User can register with email/password- [ ] User can login with valid credentials- [ ] Invalid credentials rejected with proper error- [ ] Sessions persist across page reloads- [ ] Logout clears session
Regression Evals:- [ ] Public routes still accessible- [ ] API responses unchanged- [ ] Database schema compatible
### Phase 2: Implement (varies)[Write code]
### Phase 3: EvaluateRun: /eval check add-authentication
### Phase 4: ReportEVAL REPORT: add-authentication==============================Capability: 5/5 passed (pass@3: 100%)Regression: 3/3 passed (pass^3: 100%)Status: SHIP IT

Local Framework Utilities

The mechanical utilities ship in scripts/lib/eval-harness/:

sh
node scripts/eval-harness.js example
  • Capsule: hash-linked journal with five lineages and local integrity checks.
  • Inspection: source digests, validated variant paths, and syntactic warnings.
  • Replay: declared tools and content-addressed fixtures. Missing fixtures fail closed; SE3 and above are refused in replay. Record mode invokes the registered implementation, so only register trusted functions.
  • Receipt: offline verification of capsule and artifact bytes, with named checks.
  • Retrospective preparation: node scripts/eval-harness.js capsule group <dir> [<dir> ...] groups 1 to 100 explicitly selected, verified local capsule snapshots from one task family by declared harness version. Repeated snapshots count once; conflicting identities or invalid capsules reject the whole report. This is read-only record counting, with no new rollouts, scores or promotion. Use small, quiescent capsules. Payloads, directory arguments and raw run/capsule IDs are omitted, but task-family/version labels are verbatim and digest references are linkable; review them before sharing. Operational validation remains pending.

Candidate execution is disabled on every OS because no verified OS containment backend is implemented. gate run, runGate, runVariant, direct child launch, and the retired effect preload refuse with gate.isolation_required. No trust flag or caller-supplied executor can bypass the refusal. The example records that refusal and inspects source without executing or scoring it.

Do not present static warnings, a capsule receipt, or successful utility tests as candidate containment or promotion evidence. A future gate requires an independently reviewed OS boundary, protected checker and audit channels, and fatal baseline rejection. See docs/architecture/eval-harness-frameworks.md.

Product Evals (v1.8)

Use product evals when behavior quality cannot be captured by unit tests alone.

Grader Types

  1. Code grader (deterministic assertions)
  2. Rule grader (regex/schema constraints)
  3. Model grader (LLM-as-judge rubric)
  4. Human grader (manual adjudication for ambiguous outputs)

pass@k Guidance

  • pass@1: direct reliability
  • pass@3: practical reliability under controlled retries
  • pass^3: stability test (all 3 runs must pass)

Recommended thresholds:

  • Capability evals: pass@3 >= 0.90
  • Regression evals: pass^3 = 1.00 for release-critical paths

Eval Anti-Patterns

  • Overfitting prompts to known eval examples
  • Measuring only happy-path outputs
  • Ignoring cost and latency drift while chasing pass rates
  • Allowing flaky graders in release gates

Minimal Eval Artifact Layout

  • .claude/evals/<feature>.md definition
  • .claude/evals/<feature>.log run history
  • docs/releases/<version>/eval-summary.md release snapshot

来源与署名

来源:affaan-m/ECC位于skills/eval-harness提交ef648e0

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架