Debug My Harness

zernie/vigiles/skills/debug-my-harness

by zernie1d562c69566fNo license15 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Diagnose why an agent harness misbehaved by reading the local flight-recorder ledger (.vigiles/runs.jsonl) — which skills fired or got hijacked, which hooks blocked or wrongly allowed, which subagent tool-contract violations happened, and how a skill's trigger rate moved. Use when asked why a skill stopped firing, why a hook didn't block, why the wrong skill ran, or to debug/investigate what the harness actually did. NOT for writing new rules (use strengthen) or editing the spec (use edit-spec).

Instructions onlyAI & Agents
AI-generated overview

Diagnoses agent harness misbehavior by reading the local .vigiles/runs.jsonl flight-recorder ledger.

What it does
Reads the local append-only ledger at .vigiles/runs.jsonl, which records hook decisions, subagent tool-contract decisions, skill activations, eval metrics and capability diffs. It matches a user's question to the relevant records to explain why a skill stopped firing, why a hook did not block, whether a subagent went outside its contract, or whether metrics are drifting. It then recommends a fix tied to specific records and offers handoffs to related skills. It only reads the ledger and does not write new rules or edit specs.
When to use it
Use when asked why a skill stopped firing, why the wrong skill ran, why a hook failed to block, or whether a subagent misbehaved. Also useful for investigating what the harness actually did and whether measured metrics are trending worse.
Requirements
Requires access to the local ledger file .vigiles/runs.jsonl produced by the vigiles harness; if it is absent or empty the skill reports that nothing is recorded. Uses Read, Glob and Grep tools. Ships no scripts.

Diagnose harness misbehavior from the flight recorder — the local, append-only ledger at .vigiles/runs.jsonl that vigiles writes as your harness runs. It records what actually happened, so you debug from evidence instead of guessing.

What's in the ledger

One JSON record per line, each with a kind:

  • hook — a compiled-hook gate decision: {event, decision: allow|deny|ask, mode: enforce|observe, rule, cmd, reason}.
  • agent — a subagent tool-contract decision: {name, tool, allowed, reason} (a false = the agent went outside its lane).
  • skill — a skill activation: {name, fired}.
  • eval — a measured metric: {name, metric, value} (e.g. trigger-rate recall/precision).
  • capability-diff — a blast-radius change: {pr, added, removed, widened}.

Instructions

Step 1: Read the ledger

Read .vigiles/runs.jsonl (JSONL — one record per line; tolerate a torn last line). If it's absent or empty, say so — there's nothing recorded yet; suggest running the harness (or vigiles audit) first. Do NOT fabricate records.

Step 2: Answer the specific question, evidence-first

Match the user's question to the ledger:

  • "Why did skill X stop firing / why does the wrong one run?" — count skill fires by name over time. If X's fire-rate dropped, look for a sibling that fired on the same kinds of prompts (a selection collision) and check their descriptions for overlap. Recommend differentiating or merging the descriptions.
  • "Why didn't my hook block that?" — find hook records for the event. A decision: allow on something that should be denied, or mode: observe (shadow, never blocks), or the absence of any record, tells you which. Recommend flipping observe→enforce or fixing the gate logic.
  • "Did a subagent misbehave?" — list agent records with allowed: false: the agent reached for a tool outside its declared contract. Point at the contract to tighten or widen.
  • "Is it getting worse?" — compare eval metric values (recall/precision) across runs; a downward trend is drift (often after a harness/model upgrade).

Step 3: Recommend a fix, tied to the evidence

Prefer promoting an ignored-but-decidable rule from prose to a deterministic gate: a repeated agent violation or a rule the agent keeps breaking → a compiled hook or a tighter tool-contract (the strengthen skill can help). A description collision → differentiate the skill descriptions. Always cite the specific records you based the diagnosis on.

Step 4: Offer the next step

If the fix is a spec change, hand off to edit-spec. If it's promoting guidance to a linter rule, hand off to strengthen. If a behavioral claim needs measuring (does the skill fire now?), hand off to test-harness (measureTriggerRate).

Source and attribution

Source:zernie/vigilesinskills/debug-my-harnessat commit1d562c6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from zernie/vigiles

Linter Docs

zernie

Deep linter reference for authoring or debugging a vigiles enforce() rule — plugin tables, AST selectors, type-aware rules, auto-fix, and edge cases for ESLint, Ruff, Pylint, RuboCop, Stylelint, and Clippy. Use when you need the exact rule name or config for a specific linter, not for running a linter. (JVM/Go linters — detekt, ktlint, Checkstyle, golangci-lint — and Cedar have no deep-dive file yet; their reference lives in docs/linter-support.md.)

Awaiting classification15updated today

Review Docs

zernie

Review the README or a documentation page from several real reader points of view at once. Fans out one parallel reviewer per audience — Claude Code newcomer, power user, plugin/skill author, skeptical senior engineer, non-running decision-maker — each scoring it x/5 and naming concrete, line-level fixes. Use when asked to review, critique, grade, or "get to N/5" the README or any front-door doc, or to check how a doc reads for real users. Not for code review (use /code-review for that).

Awaiting classification15updated today

Adopt Spec

zernie

Converts an existing hand-written CLAUDE.md into a typed CLAUDE.md.spec.ts using the vigiles spec library.

AI & Agents15updated today

Pr To Lint Rule

zernie

Turns a recurring prose code-review rule into a custom lint rule, gated by an independent soundness test that abstains if it leaks.

Software Development15updated today

Enforce Rules Format

zernie

Validates that project instruction-file rules carry proper enforcement classifications and fixes missing ones.

AI & Agents15updated today

Deep Research

zernie

Use when the user asks to research a topic in depth, map a competitive/market landscape, run a multi-source investigation, or "fan out" parallel research agents — anything where many findings must be gathered and then NOT lost. Enforces durable, detail-preserving research (write full findings to disk; keep a full appendix beside the synthesis).

Awaiting classification15updated today
Debug My Harness · skills/debug-my-harness Agent Skill | SourceWeft