Agent Eval

affaan-m/ECC/skills/agent-eval

作者 affaan-mef648e01899ba3e8dc6371642deaaf64b4477775MIT275K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫4 天前更新

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression.

僅含說明AI & Agents
AI 產生的概覽

在自訂任務上對程式編寫代理進行對照評測,輸出通過率、成本、時間與一致性指標。

功能
此技能說明一套命令列工作流程,用於在可重現的任務上對 Claude Code、Aider、Codex 等程式編寫代理進行基準比較。任務以 YAML 宣告,包含提示詞、涉及檔案、固定提交與評判標準,每次代理執行都在獨立的 git worktree 中隔離進行。它會記錄通過率、API 成本、實際耗時以及多次重複執行的一致性,並產生比較報表。
適用情境
適用於在自己的程式庫上比較不同程式編寫代理或模型,或在調整代理設定後需要可量化的通過率、成本與時間,而非主觀印象時。也適合在代理更新模型或工具鏈後進行回歸檢查,以及為團隊提供有資料依據的選型決策。
執行需求
需要從程式碼倉庫安裝 agent-eval 命令列工具、用於 worktree 隔離的 git,以及被比較的各個代理。評判標準可能依賴 pytest、npm 等測試或建置工具,以模型為基礎的評判需要模型存取權限;成本指標取決於能否取得 API 花費資料。

Agent Eval Skill

A lightweight CLI tool for comparing coding agents head-to-head on reproducible tasks. Every "which coding agent is best?" comparison runs on vibes — this tool systematizes it.

When to Activate

  • Comparing coding agents (Claude Code, Aider, Codex, etc.) on your own codebase
  • Measuring agent performance before adopting a new tool or model
  • Running regression checks when an agent updates its model or tooling
  • Producing data-backed agent selection decisions for a team

Installation

Note: Install agent-eval from its repository after reviewing the source.

Core Concepts

YAML Task Definitions

Define tasks declaratively. Each task specifies what to do, which files to touch, and how to judge success:

yaml
name: add-retry-logicdescription: Add exponential backoff retry to the HTTP clientrepo: ./my-projectfiles:  - src/http_client.pyprompt: |  Add retry logic with exponential backoff to all HTTP requests.  Max 3 retries. Initial delay 1s, max delay 30s.judge:  - type: pytest    command: pytest tests/test_http_client.py -v  - type: grep    pattern: "exponential_backoff|retry"    files: src/http_client.pycommit: "abc1234"  # pin to specific commit for reproducibility

Git Worktree Isolation

Each agent run gets its own git worktree — no Docker required. This provides reproducibility isolation so agents cannot interfere with each other or corrupt the base repo.

Metrics Collected

MetricWhat It Measures
Pass rateDid the agent produce code that passes the judge?
CostAPI spend per task (when available)
TimeWall-clock seconds to completion
ConsistencyPass rate across repeated runs (e.g., 3/3 = 100%)

Workflow

1. Define Tasks

Create a tasks/ directory with YAML files, one per task:

bash
mkdir tasks# Write task definitions (see template above)

2. Run Agents

Execute agents against your tasks:

bash
agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3

Each run:

  1. Creates a fresh git worktree from the specified commit
  2. Hands the prompt to the agent
  3. Runs the judge criteria
  4. Records pass/fail, cost, and time

3. Compare Results

Generate a comparison report:

bash
agent-eval report --format table
Task: add-retry-logic (3 runs each)┌──────────────┬───────────┬────────┬────────┬─────────────┐│ Agent        │ Pass Rate │ Cost   │ Time   │ Consistency │├──────────────┼───────────┼────────┼────────┼─────────────┤│ claude-code  │ 3/3       │ $0.12  │ 45s    │ 100%        ││ aider        │ 2/3       │ $0.08  │ 38s    │  67%        │└──────────────┴───────────┴────────┴────────┴─────────────┘

Judge Types

Code-Based (deterministic)

yaml
judge:  - type: pytest    command: pytest tests/ -v  - type: command    command: npm run build

Pattern-Based

yaml
judge:  - type: grep    pattern: "class.*Retry"    files: src/**/*.py

Model-Based (LLM-as-judge)

yaml
judge:  - type: llm    prompt: |      Does this implementation correctly handle exponential backoff?      Check for: max retries, increasing delays, jitter.

Best Practices

  • Start with 3-5 tasks that represent your real workload, not toy examples
  • Run at least 3 trials per agent to capture variance — agents are non-deterministic
  • Pin the commit in your task YAML so results are reproducible across days/weeks
  • Include at least one deterministic judge (tests, build) per task — LLM judges add noise
  • Track cost alongside pass rate — a 95% agent at 10x the cost may not be the right choice
  • Version your task definitions — they are test fixtures, treat them as code

Links

來源與署名

來源:affaan-m/ECC位於skills/agent-eval提交ef648e0

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架