Compares coding agents head-to-head on reproducible YAML tasks, measuring pass rate, cost, time and consistency.
- What it does
- This skill documents a lightweight CLI workflow for benchmarking coding agents such as Claude Code, Aider and Codex on the same tasks. Tasks are declared in YAML files that specify the prompt, target files, a pinned commit and judge criteria, and each agent run gets its own git worktree for isolation. It records pass/fail, cost, wall-clock time and consistency across repeated runs, then produces a comparison report table.
- When to use it
- Use it when choosing between coding agents or models for your own codebase, before adopting a new tool, or when re-checking agent performance after a model or tool update. It also suits teams that want data-backed agent selection decisions.
- Requirements
- Requires the agent-eval CLI installed from its repository, git for worktree isolation, and the agents being compared. Judge types may need pytest, npm or similar commands, and model-based judging needs an LLM. No scripts ship with the skill; it is instructions only.
エージェント評価スキル
再現可能なタスクでコーディングエージェントをヘッドツーヘッドで比較するための軽量 CLI ツールです。「どのコーディングエージェントが最適か?」という比較はすべて感覚に頼りがちです — このツールはそれを体系化します。
起動タイミング
- 自分のコードベースでコーディングエージェント(Claude Code、Aider、Codex など)を比較する
- 新しいツールやモデルを採用する前にエージェントパフォーマンスを測定する
- エージェントがモデルやツールを更新した際にリグレッションチェックを実行する
- チームにデータに基づいたエージェント選択の判断を提供する
インストール
注意: agent-eval はソースを確認した後、リポジトリからインストールしてください。
コアコンセプト
YAML タスク定義
タスクを宣言的に定義します。各タスクは何をするか、どのファイルを操作するか、成功をどう判定するかを指定します:
Git ワークツリー分離
各エージェント実行は独自の git ワークツリーを取得します — Docker 不要。これにより再現性の分離が提供され、エージェントが互いに干渉したりベースリポジトリを破壊したりしません。
収集メトリクス
ワークフロー
1. タスクの定義
タスクごとに 1 つの YAML ファイルを持つ tasks/ ディレクトリを作成します:
2. エージェントの実行
タスクに対してエージェントを実行します:
各実行:
- 指定されたコミットから新しい git ワークツリーを作成
- エージェントにプロンプトを渡す
- ジャッジ基準を実行
- 合格・不合格、コスト、時間を記録
3. 結果の比較
比較レポートを生成します:
ジャッジタイプ
コードベース(決定論的)
パターンベース
モデルベース(LLM-as-judge)
ベストプラクティス
- 3〜5 タスクから始める — おもちゃの例ではなく、実際のワークロードを代表するタスク
- エージェントごとに少なくとも 3 試行実行する — エージェントは非決定論的なので分散を把握する
- タスク YAML でコミットを固定する — 日や週をまたいで結果が再現可能になる
- タスクごとに少なくとも 1 つの決定論的ジャッジを含める(テスト、ビルド)— LLM ジャッジはノイズを加える
- 合格率と一緒にコストを追跡する — 10 倍のコストで 95% のエージェントが正しい選択でない場合もある
- タスク定義をバージョン管理する — それらはテストフィクスチャであり、コードとして扱う
リンク