Agent Eval

colbymchenry/codegraph/.claude/skills/agent-eval

作者 colbymchenryb635dd467f05無授權條款73K 個星標收錄於 2026年10月9日更新於 2026年10月8日儲存庫昨天更新

Benchmark CodeGraph retrieval quality on a real codebase by comparing agent behavior with vs without CodeGraph. Use when the user runs /agent-eval or asks to test, benchmark, audit, or validate a codegraph version (the local dev build or a published npm version) against a language's repo.

僅含說明AI & Agents
AI 產生的概覽

透過在真實程式碼庫上比較有無 CodeGraph 時的代理行為,對 CodeGraph 檢索品質進行基準測試。

功能
此技能執行品質稽核,衡量在選定的 CodeGraph 版本與選定的真實程式碼庫上,CodeGraph 相較於單純 grep 與檔案讀取能為代理帶來多少助益。它引導使用者依序選擇版本、語言、程式碼庫與測試模式,接著啟動背景稽核指令碼,該指令碼會複製並重新建立程式碼庫索引,並執行對照測試組。它會報告各測試組的指標,例如工具呼叫次數、檔案讀取次數、耗時、成本與回饋指標,並提供並排比較表。
適用情境
當使用者執行 /agent-eval,或要求針對某語言的程式碼庫測試、基準測試、稽核或驗證某個 CodeGraph 版本(本機開發組建或已發佈的 npm 版本)時使用。
執行需求
需要 tmux 3+、已登入的 claude CLI、node 與 git,執行於 macOS 或 Linux,並從 CodeGraph 儲存庫根目錄執行。它會讀取 corpus.json 檔案,驅動 scripts/agent-eval/ 下的測試指令碼,將語料庫儲存庫複製到 /tmp,並在測試期間暫時修改全域 codegraph 安裝後再還原。

CodeGraph Quality Audit

Measures how much CodeGraph helps an agent versus plain grep/read, for a chosen codegraph version on a chosen real-world repo. Drives the harness in scripts/agent-eval/.

Prerequisites

  • tmux 3+, a logged-in claude CLI, node, git (macOS/Linux).
  • Run from the codegraph repo root.

Workflow

Copy this checklist:

- [ ] 1. Pick version (local or npm)- [ ] 2. Pick language- [ ] 3. Pick repo by size- [ ] 4. Pick harness (headless / tmux / both)- [ ] 5. Run audit.sh in the background- [ ] 6. Report results

Step 1 — version. Ask with AskUserQuestion: which codegraph version to test. Offer "Local dev build" and "Latest published"; the free-text "Other" lets the user type a specific version (e.g. 0.7.10). Map the answer to a VERSION token:

  • "Local dev build" → local
  • "Latest published" → latest
  • a typed version → that string (e.g. 0.7.10)

Step 2 — language. Read .claude/skills/agent-eval/corpus.json. Ask with AskUserQuestion which language to test, listing the languages that have entries.

Step 3 — repo. From the chosen language's entries, ask which repo. Label each option with its size and file count, e.g. excalidraw — Medium (~600 files). Each entry carries the repo URL and a representative question.

Step 4 — harness. Ask with AskUserQuestion which harness to run, and map the answer to a MODE token:

  • "Headless" → headless — claude -p with stream-json: exact tokens/cost and a clean tool sequence (2 runs, fast, no TTY).
  • "Interactive (tmux)" → tmux — drives the real Claude TUI in tmux: faithful Explore-subagent behavior, metrics from session logs (2 runs, slower).
  • "Both" → all — headless + interactive (4 runs).

Step 5 — run. Launch in the background (sets the version, clones if missing, wipes + re-indexes, runs the chosen arms — several minutes):

bash
scripts/agent-eval/audit.sh <VERSION> <repo-name> <repo-url> "<question>" <MODE>

Step 6 — report. When the job finishes, read the log and report per arm:

  • Headless (parse-run.mjs): total tool calls, file Reads, Grep/Bash, codegraph-tool calls, duration, total cost.
  • Interactive (parse-session.mjs): the VERDICT: codegraph_explore used Nx | Read N | Grep/Bash N and TOKENS: lines.
  • Both paths also print the three feedback metrics — residual context occupancy, explore sufficiency, allocation efficiency — and a headless A/B ends with a side-by-side ARM COMPARISON table. Report that table, and check its contamination row first: CLI calls that RETURNED output > 0 means the arm reached codegraph through Bash and its numbers are void. How to read the rest: docs/benchmarks/agent-eval-feedback-metrics.md.

Lead with cost + tool/Read counts — they are the reliable signals; raw token in/out are confounded by subagent delegation and prompt caching. State whether codegraph reduced effort and whether both arms reached a correct answer.

Notes

  • The index is rebuilt every run (audit.sh wipes .codegraph) — different versions extract differently, so an index must be served by the same binary that built it.
  • audit.sh temporarily mutates the global codegraph install for the test, then restores your dev link via local-install.sh.
  • Corpus repos are cloned to /tmp/codegraph-corpus (reused if already present).
  • Add or edit repos in corpus.json (fields: name, repo, size, files, question).

來源與署名

來源:colbymchenry/codegraph位於.claude/skills/agent-eval提交b635dd4

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架