Agent Eval

affaan-m/ECC/pi/core/skills/agent-eval

作者 affaan-mef648e01899ba3e8dc6371642deaaf64b4477775MIT275K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库4天前更新

Head-to-head comparison of coding agents (Claude Code, Aider, Codex, etc.) on custom tasks with pass rate, cost, time, and consistency metrics. Use when choosing between coding agents, or when a change to an agent setup needs measured pass rate, cost, and time rather than an impression.

仅含说明AI & Agents
AI 生成的概览

在自定义任务上对编程智能体进行对比评测,输出通过率、成本、耗时与一致性指标。

功能
该技能描述了一套命令行工作流,用于在可复现任务上对 Claude Code、Aider、Codex 等编程智能体进行基准对比。任务以 YAML 声明,包含提示词、涉及文件、固定提交和评判标准,每次智能体运行都在独立的 git worktree 中隔离执行。它会记录通过率、API 成本、实际耗时以及多次重复运行的一致性,并生成对比报告表格。
适用场景
适用于在自己的代码库上比较不同编程智能体或模型,或在调整智能体配置后需要可量化的通过率、成本与耗时而非主观印象时。也适合在智能体更新模型或工具链后做回归检查,以及为团队提供有数据支撑的选型决策。
运行要求
需要从代码仓库安装 agent-eval 命令行工具、用于 worktree 隔离的 git,以及被比较的各个智能体。评判标准可能依赖 pytest、npm 等测试或构建工具,基于模型的评判需要模型访问权限;成本指标取决于能否获取 API 花费数据。

Agent Eval Skill

A lightweight CLI tool for comparing coding agents head-to-head on reproducible tasks. Every "which coding agent is best?" comparison runs on vibes — this tool systematizes it.

When to Activate

  • Comparing coding agents (Claude Code, Aider, Codex, etc.) on your own codebase
  • Measuring agent performance before adopting a new tool or model
  • Running regression checks when an agent updates its model or tooling
  • Producing data-backed agent selection decisions for a team

Installation

Note: Install agent-eval from its repository after reviewing the source.

Core Concepts

YAML Task Definitions

Define tasks declaratively. Each task specifies what to do, which files to touch, and how to judge success:

yaml
name: add-retry-logicdescription: Add exponential backoff retry to the HTTP clientrepo: ./my-projectfiles:  - src/http_client.pyprompt: |  Add retry logic with exponential backoff to all HTTP requests.  Max 3 retries. Initial delay 1s, max delay 30s.judge:  - type: pytest    command: pytest tests/test_http_client.py -v  - type: grep    pattern: "exponential_backoff|retry"    files: src/http_client.pycommit: "abc1234"  # pin to specific commit for reproducibility

Git Worktree Isolation

Each agent run gets its own git worktree — no Docker required. This provides reproducibility isolation so agents cannot interfere with each other or corrupt the base repo.

Metrics Collected

MetricWhat It Measures
Pass rateDid the agent produce code that passes the judge?
CostAPI spend per task (when available)
TimeWall-clock seconds to completion
ConsistencyPass rate across repeated runs (e.g., 3/3 = 100%)

Workflow

1. Define Tasks

Create a tasks/ directory with YAML files, one per task:

bash
mkdir tasks# Write task definitions (see template above)

2. Run Agents

Execute agents against your tasks:

bash
agent-eval run --task tasks/add-retry-logic.yaml --agent claude-code --agent aider --runs 3

Each run:

  1. Creates a fresh git worktree from the specified commit
  2. Hands the prompt to the agent
  3. Runs the judge criteria
  4. Records pass/fail, cost, and time

3. Compare Results

Generate a comparison report:

bash
agent-eval report --format table
Task: add-retry-logic (3 runs each)┌──────────────┬───────────┬────────┬────────┬─────────────┐│ Agent        │ Pass Rate │ Cost   │ Time   │ Consistency │├──────────────┼───────────┼────────┼────────┼─────────────┤│ claude-code  │ 3/3       │ $0.12  │ 45s    │ 100%        ││ aider        │ 2/3       │ $0.08  │ 38s    │  67%        │└──────────────┴───────────┴────────┴────────┴─────────────┘

Judge Types

Code-Based (deterministic)

yaml
judge:  - type: pytest    command: pytest tests/ -v  - type: command    command: npm run build

Pattern-Based

yaml
judge:  - type: grep    pattern: "class.*Retry"    files: src/**/*.py

Model-Based (LLM-as-judge)

yaml
judge:  - type: llm    prompt: |      Does this implementation correctly handle exponential backoff?      Check for: max retries, increasing delays, jitter.

Best Practices

  • Start with 3-5 tasks that represent your real workload, not toy examples
  • Run at least 3 trials per agent to capture variance — agents are non-deterministic
  • Pin the commit in your task YAML so results are reproducible across days/weeks
  • Include at least one deterministic judge (tests, build) per task — LLM judges add noise
  • Track cost alongside pass rate — a 95% agent at 10x the cost may not be the right choice
  • Version your task definitions — they are test fixtures, treat them as code

Links

来源与署名

来源:affaan-m/ECC位于pi/core/skills/agent-eval提交ef648e0

许可证: MIT

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架