Debug My Harness

zernie/vigiles/skills/debug-my-harness

作者 zernie1d562c69566f无许可证15 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Diagnose why an agent harness misbehaved by reading the local flight-recorder ledger (.vigiles/runs.jsonl) — which skills fired or got hijacked, which hooks blocked or wrongly allowed, which subagent tool-contract violations happened, and how a skill's trigger rate moved. Use when asked why a skill stopped firing, why a hook didn't block, why the wrong skill ran, or to debug/investigate what the harness actually did. NOT for writing new rules (use strengthen) or editing the spec (use edit-spec).

仅含说明AI & Agents
AI 生成的概览

通过读取本地 .vigiles/runs.jsonl 飞行记录账本来诊断智能体框架的异常行为。

功能
读取本地只追加账本 .vigiles/runs.jsonl,其中记录钩子决策、子智能体工具契约决策、技能激活、评估指标和能力差异。它把用户的问题与相关记录对应起来,解释技能为何不再触发、钩子为何没有拦截、子智能体是否越出契约,或指标是否在漂移。随后给出与具体记录挂钩的修复建议,并提供移交给相关技能的下一步。它只读取账本,不编写新规则,也不编辑规范。
适用场景
适用于询问技能为何停止触发、为何运行了错误的技能、钩子为何未能拦截,或子智能体是否行为不当。也适合调查框架实际做了什么,以及测量指标是否在变差。
运行要求
需要访问由 vigiles 框架生成的本地账本文件 .vigiles/runs.jsonl;若文件缺失或为空,该技能会说明尚无记录。使用 Read、Glob 和 Grep 工具。不附带脚本。

Diagnose harness misbehavior from the flight recorder — the local, append-only ledger at .vigiles/runs.jsonl that vigiles writes as your harness runs. It records what actually happened, so you debug from evidence instead of guessing.

What's in the ledger

One JSON record per line, each with a kind:

  • hook — a compiled-hook gate decision: {event, decision: allow|deny|ask, mode: enforce|observe, rule, cmd, reason}.
  • agent — a subagent tool-contract decision: {name, tool, allowed, reason} (a false = the agent went outside its lane).
  • skill — a skill activation: {name, fired}.
  • eval — a measured metric: {name, metric, value} (e.g. trigger-rate recall/precision).
  • capability-diff — a blast-radius change: {pr, added, removed, widened}.

Instructions

Step 1: Read the ledger

Read .vigiles/runs.jsonl (JSONL — one record per line; tolerate a torn last line). If it's absent or empty, say so — there's nothing recorded yet; suggest running the harness (or vigiles audit) first. Do NOT fabricate records.

Step 2: Answer the specific question, evidence-first

Match the user's question to the ledger:

  • "Why did skill X stop firing / why does the wrong one run?" — count skill fires by name over time. If X's fire-rate dropped, look for a sibling that fired on the same kinds of prompts (a selection collision) and check their descriptions for overlap. Recommend differentiating or merging the descriptions.
  • "Why didn't my hook block that?" — find hook records for the event. A decision: allow on something that should be denied, or mode: observe (shadow, never blocks), or the absence of any record, tells you which. Recommend flipping observe→enforce or fixing the gate logic.
  • "Did a subagent misbehave?" — list agent records with allowed: false: the agent reached for a tool outside its declared contract. Point at the contract to tighten or widen.
  • "Is it getting worse?" — compare eval metric values (recall/precision) across runs; a downward trend is drift (often after a harness/model upgrade).

Step 3: Recommend a fix, tied to the evidence

Prefer promoting an ignored-but-decidable rule from prose to a deterministic gate: a repeated agent violation or a rule the agent keeps breaking → a compiled hook or a tighter tool-contract (the strengthen skill can help). A description collision → differentiate the skill descriptions. Always cite the specific records you based the diagnosis on.

Step 4: Offer the next step

If the fix is a spec change, hand off to edit-spec. If it's promoting guidance to a linter rule, hand off to strengthen. If a behavioral claim needs measuring (does the skill fire now?), hand off to test-harness (measureTriggerRate).

来源与署名

来源:zernie/vigiles位于skills/debug-my-harness提交1d562c6

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 zernie/vigiles 的技能

Linter Docs

zernie

Deep linter reference for authoring or debugging a vigiles enforce() rule — plugin tables, AST selectors, type-aware rules, auto-fix, and edge cases for ESLint, Ruff, Pylint, RuboCop, Stylelint, and Clippy. Use when you need the exact rule name or config for a specific linter, not for running a linter. (JVM/Go linters — detekt, ktlint, Checkstyle, golangci-lint — and Cedar have no deep-dive file yet; their reference lives in docs/linter-support.md.)

待分类15今天更新

Review Docs

zernie

Review the README or a documentation page from several real reader points of view at once. Fans out one parallel reviewer per audience — Claude Code newcomer, power user, plugin/skill author, skeptical senior engineer, non-running decision-maker — each scoring it x/5 and naming concrete, line-level fixes. Use when asked to review, critique, grade, or "get to N/5" the README or any front-door doc, or to check how a doc reads for real users. Not for code review (use /code-review for that).

待分类15今天更新

Adopt Spec

zernie

将现有的手写 CLAUDE.md 转换为使用 vigiles 规范库的类型化 CLAUDE.md.spec.ts。

AI & Agents15今天更新

Pr To Lint Rule

zernie

把反复出现的散文式代码审查规则合成为自定义 lint 规则,并由独立健全性测试把关,不通过则弃用。

Software Development15今天更新

Enforce Rules Format

zernie

验证项目指令文件中的规则是否具有正确的执行分类,并修复缺失的分类。

AI & Agents15今天更新

Deep Research

zernie

Use when the user asks to research a topic in depth, map a competitive/market landscape, run a multi-source investigation, or "fan out" parallel research agents — anything where many findings must be gathered and then NOT lost. Enforces durable, detail-preserving research (write full findings to disk; keep a full appendix beside the synthesis).

待分类15今天更新