Debug My Harness

zernie/vigiles/skills/debug-my-harness

作者 zernie1d562c69566f無授權條款15 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Diagnose why an agent harness misbehaved by reading the local flight-recorder ledger (.vigiles/runs.jsonl) — which skills fired or got hijacked, which hooks blocked or wrongly allowed, which subagent tool-contract violations happened, and how a skill's trigger rate moved. Use when asked why a skill stopped firing, why a hook didn't block, why the wrong skill ran, or to debug/investigate what the harness actually did. NOT for writing new rules (use strengthen) or editing the spec (use edit-spec).

僅含說明AI & Agents
AI 產生的概覽

透過讀取本機 .vigiles/runs.jsonl 飛行記錄帳本來診斷代理框架的異常行為。

功能
讀取本機僅附加帳本 .vigiles/runs.jsonl,其中記錄掛鉤決策、子代理工具契約決策、技能啟用、評估指標與能力差異。它會把使用者的問題對應到相關紀錄,解釋技能為何不再觸發、掛鉤為何沒有攔阻、子代理是否超出契約,或指標是否正在漂移。接著提出與具體紀錄綁定的修正建議,並提供移交給相關技能的下一步。它只讀取帳本,不撰寫新規則,也不編輯規格。
適用情境
適用於詢問技能為何停止觸發、為何執行到錯誤的技能、掛鉤為何未能攔阻,或子代理是否行為不當。也適合調查框架實際做了什麼,以及量測指標是否正在變差。
執行需求
需要存取由 vigiles 框架產生的本機帳本檔案 .vigiles/runs.jsonl;若檔案不存在或為空,該技能會說明尚無紀錄。使用 Read、Glob 與 Grep 工具。不附帶指令碼。

Diagnose harness misbehavior from the flight recorder — the local, append-only ledger at .vigiles/runs.jsonl that vigiles writes as your harness runs. It records what actually happened, so you debug from evidence instead of guessing.

What's in the ledger

One JSON record per line, each with a kind:

  • hook — a compiled-hook gate decision: {event, decision: allow|deny|ask, mode: enforce|observe, rule, cmd, reason}.
  • agent — a subagent tool-contract decision: {name, tool, allowed, reason} (a false = the agent went outside its lane).
  • skill — a skill activation: {name, fired}.
  • eval — a measured metric: {name, metric, value} (e.g. trigger-rate recall/precision).
  • capability-diff — a blast-radius change: {pr, added, removed, widened}.

Instructions

Step 1: Read the ledger

Read .vigiles/runs.jsonl (JSONL — one record per line; tolerate a torn last line). If it's absent or empty, say so — there's nothing recorded yet; suggest running the harness (or vigiles audit) first. Do NOT fabricate records.

Step 2: Answer the specific question, evidence-first

Match the user's question to the ledger:

  • "Why did skill X stop firing / why does the wrong one run?" — count skill fires by name over time. If X's fire-rate dropped, look for a sibling that fired on the same kinds of prompts (a selection collision) and check their descriptions for overlap. Recommend differentiating or merging the descriptions.
  • "Why didn't my hook block that?" — find hook records for the event. A decision: allow on something that should be denied, or mode: observe (shadow, never blocks), or the absence of any record, tells you which. Recommend flipping observe→enforce or fixing the gate logic.
  • "Did a subagent misbehave?" — list agent records with allowed: false: the agent reached for a tool outside its declared contract. Point at the contract to tighten or widen.
  • "Is it getting worse?" — compare eval metric values (recall/precision) across runs; a downward trend is drift (often after a harness/model upgrade).

Step 3: Recommend a fix, tied to the evidence

Prefer promoting an ignored-but-decidable rule from prose to a deterministic gate: a repeated agent violation or a rule the agent keeps breaking → a compiled hook or a tighter tool-contract (the strengthen skill can help). A description collision → differentiate the skill descriptions. Always cite the specific records you based the diagnosis on.

Step 4: Offer the next step

If the fix is a spec change, hand off to edit-spec. If it's promoting guidance to a linter rule, hand off to strengthen. If a behavioral claim needs measuring (does the skill fire now?), hand off to test-harness (measureTriggerRate).

來源與署名

來源:zernie/vigiles位於skills/debug-my-harness提交1d562c6

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 zernie/vigiles 的技能

Linter Docs

zernie

Deep linter reference for authoring or debugging a vigiles enforce() rule — plugin tables, AST selectors, type-aware rules, auto-fix, and edge cases for ESLint, Ruff, Pylint, RuboCop, Stylelint, and Clippy. Use when you need the exact rule name or config for a specific linter, not for running a linter. (JVM/Go linters — detekt, ktlint, Checkstyle, golangci-lint — and Cedar have no deep-dive file yet; their reference lives in docs/linter-support.md.)

待分類15今天更新

Review Docs

zernie

Review the README or a documentation page from several real reader points of view at once. Fans out one parallel reviewer per audience — Claude Code newcomer, power user, plugin/skill author, skeptical senior engineer, non-running decision-maker — each scoring it x/5 and naming concrete, line-level fixes. Use when asked to review, critique, grade, or "get to N/5" the README or any front-door doc, or to check how a doc reads for real users. Not for code review (use /code-review for that).

待分類15今天更新

Adopt Spec

zernie

將現有的手寫 CLAUDE.md 轉換為使用 vigiles 規範函式庫的型別化 CLAUDE.md.spec.ts。

AI & Agents15今天更新

Pr To Lint Rule

zernie

把反覆出現的散文式程式碼審查規則合成為自訂 lint 規則,並由獨立健全性測試把關,未通過就棄用。

Software Development15今天更新

Enforce Rules Format

zernie

驗證專案指令檔案中的規則是否具備正確的強制分類,並修正缺少的分類。

AI & Agents15今天更新

Deep Research

zernie

Use when the user asks to research a topic in depth, map a competitive/market landscape, run a multi-source investigation, or "fan out" parallel research agents — anything where many findings must be gathered and then NOT lost. Enforces durable, detail-preserving research (write full findings to disk; keep a full appendix beside the synthesis).

待分類15今天更新