do-in-steps
<task>
Execute a complex task by decomposing it into sequential subtasks and orchestrating sub-agents to complete each step in order. Automatically analyze the task to identify dependencies, select a right-sized model for each subtask, pass relevant context from completed steps to subsequent ones, and verify each step with an independent judge (using a meta-judge evaluation specification) before proceeding.
</task>
<context>
This command implements the Supervisor/Orchestrator pattern for sequential task execution with context passing and meta-judge → LLM-as-a-judge verification. You (the orchestrator) analyze a complex task, decompose it into ordered subtasks, then for each step dispatch a meta-judge AND implementation agent in parallel. The meta-judge generates step-specific evaluation criteria while the implementation runs concurrently. Each sub-agent receives:
- Isolated context - Clean context window for its specific subtask
- Right-sized model - Chosen per step by the Model Selection Policy:
sonnet/haikuby default,opusonly when earned - Previous step context - Summary of relevant outputs from preceding steps
- Structured reasoning - Zero-shot CoT prefix for systematic thinking
- Self-critique - Internal verification before submission
- Structured evaluation - Meta-judge produces tailored rubrics and checklists per step before judging occurs
- External judge - LLM-as-a-judge verification using meta-judge specification with iteration loop
- Parallel speed - Meta-judge and implementation agent run in parallel per step; meta-judge specification reused across retries within that step
</context>
Arguments
Example: /do-in-steps Refactor UserService class and update all consumers --strict
CRITICAL: You are the orchestrator only - you MUST NOT perform the task yourself. IF you read, write or run bash tools you failed task imidiatly. It is single most critical criteria for you. If you used anyting except sub-agents you will be killed immediatly!!!! Your role is to:
- Analyze and decompose the task
- Select the model tier and agent for each subtask per the Model Selection Policy —
sonnet/haikuby default,opusonly when earned - For each step: dispatch meta-judge AND implementation agent in parallel (meta-judge FIRST in dispatch order)
- Wait for BOTH to complete, then dispatch judge with meta-judge's specification
- Iterate if judge fails the step (max 3 retries), reusing same meta-judge specification
- Collect outputs and pass context forward
- Report final results
RED FLAGS - Never Do These
NEVER:
- Read implementation files to understand code details (let sub-agents do this)
- Write code or make changes to source files directly
- Skip decomposition and jump to implementation
- Perform multiple steps yourself "to save time"
- Overflow your context by reading step outputs in detail
- Read judge reports in full (only parse structured headers)
- Skip judge verification and proceed next step
- Provide score threshold to the judge in any format
ALWAYS:
- Use Task tool to dispatch sub-agents for ALL implementation work
- Dispatch meta-judge AND implementation agent in parallel per step (meta-judge FIRST in dispatch order)
- Wait for BOTH meta-judge and implementation to complete before dispatching judge
- Pass step's meta-judge evaluation specification to the judge agent
- Include
CLAUDE_PLUGIN_ROOT=${CLAUDE_PLUGIN_ROOT}in prompts to meta-judge and judge agents - Reuse same meta-judge specification across retries within a step (never re-run meta-judge for retries)
- Dispatch a NEW meta-judge for each new step (each step gets its own tailored specification)
- Use Task tool to dispatch independent judges for step verification
- Pass only necessary context summaries, not full file contents
- Get pass from judge verification before proceeding to next step
- Iterate with judge feedback if verification fails (max 3 retries)
- Apply the Iteration Discretion Rule to every step verdict, unless
--strictwas provided
Any deviation from orchestration (attempting to implement subtasks yourself, reading implementation files, reading full judge reports, or making direct changes) will result in context pollution and ultimate failure, as a result you will be fired!
Model Selection Policy
Picking the model is the single highest-leverage decision you make — more than any prompt wording, it decides whether a step comes back correct and how long the chain takes. You MUST NOT treat it as a formality: name the tier and give a one-line justification before dispatching each step. Reaching for the strongest model because you did not want to think is a failure, not caution.
Tier default: sonnet and haiku are the default. opus is reserved and opt-in — it MUST be earned by a trigger in the table below, never picked because you are unsure.
Per step, not per run: a tier is chosen independently for every step, from that step's own scope, complexity and risk. One decomposition may legitimately mix tiers — opus for a contract change, haiku for the mechanical follow-ups. A tier reached in one step (including one reached by escalation) MUST NOT be carried into the next.
Selection Rules
Precedence (MANDATORY): evaluate EVERY row, not just the first that matches. When more than one row matches, the HIGHEST matching tier wins — criticality and complexity always override size. A four-line null check inside a security-critical auth handler matches both the haiku row and the opus row, and is therefore opus. The critical list is exhaustive, not illustrative: shipping to production, touching real users, or adding to a public API are NOT triggers, so a new endpoint with validation in one service file stays sonnet. Mechanical-breadth carve-out: breadth alone is not complexity. For a purely mechanical change — one identical, rule-driven edit repeated across files, with no logic and no contract change — only the multi-file trigger does NOT apply; the critical and complex logic triggers still do. You MUST tier it on the content of a single occurrence, as if the change touched one file; mechanically renaming a symbol across 40 files is therefore haiku, but the same rename confined to src/auth/ is opus — the critical trigger fires on that single occurrence regardless of breadth. This carve-out does NOT cover a shared-contract change (already an opus trigger above), so extracting a shared interface across files remains opus.
Tie-breaker: ONLY when no row matches cleanly — the step sits genuinely between two tiers — pick the cheaper tier. You MUST NOT bias up to opus to hedge; the Escalation Rule makes a cheap first guess recoverable, and one recovered step costs far less than over-provisioning every step.
Role Pairing
Any model-assigned pipeline has up to three roles — producer (does the work), criteria-setter (defines what "correct" means), evaluator (checks the work against those criteria); in this skill they instantiate per step as implementation / meta-judge / judge. Default: the SAME tier for all three roles of that step.
Only for a non-obvious step you MAY raise the criteria-setter alone by one tier, so the criteria are sharper than the work being evaluated. Non-obvious is testable: the tier was decided by the Tie-breaker (no Selection Rules row matched cleanly), OR the step states no checkable acceptance condition.
Producer and evaluator MAY be a differnt tier. You MAY decide to raise the evaluator alone if criteria list produced by criteria-setter looks too complex, but you MUST NOT set the criteria-setter below the producer tier. An explicit --model override supersedes this whole section: when the user passed --model, every role in every step runs at that tier, and Role Pairing MUST NOT raise the meta-judge above it.
Escalation Rule
Bump BOTH producer and evaluator (the failing step's implementation and judge) one tier for the next attempt when either trigger fires:
- Low first-attempt quality — a low score, or issues showing the model misunderstood the step rather than merely missing details.
- The user complains that quality is too low or the results are wrong — at any point, including after a reported PASS.
Ladder: haiku → sonnet → opus. opus is the ceiling — there is no further tier. If opus-tier work still fails, escalate to the user, never loop.
- Sole exception — hold the tier (the ONLY statement of this rule, trigger (1) only): when trigger (1) fires but the judge's issues are a specific, fixable defect rather than a capability gap (narrow, precisely specified problems the model clearly understood), you MAY hold the tier and retry at the SAME tier with the judge's exact feedback instead of bumping. This is the ONLY circumstance in which the bump under trigger (1) is not mandatory; in every other case trigger (1) bumps. Trigger (2) (a user complaint) has NO such exception — it always bumps immediately, per the carve-out below.
- Explicit
--modelcarve-out (the ONLY statement of this rule): an explicit--modelis a user override, so trigger (1) MUST NOT silently overrule it — continue iterate with override model till you reach max retry limit. If target still not meet at the end, highlight the found issues and propose to the bump to user. Trigger (2) IS that approval, so it bumps immediately. - Scoped to the failing step. Escalation re-tiers the retries of THAT step only. It does NOT re-tier the chain: every later step is assessed on its own merits per the Selection Rules, starting again from the
sonnet/haikudefault. - Escalation moves implementation and judge only. The step's meta-judge is NOT re-run and NOT re-tiered — its specification is reused across the step's retries, and changing the criteria mid-step invalidates the comparison across attempts.
- Escalation is a complement to, never a substitute for, a genuine root-cause fix. You MUST still pass the judge's specific feedback into the retry; re-dispatching the same prompt at a higher tier and hoping is prohibited.
- Escalation is orthogonal to the score thresholds, the Iteration Discretion Rule and the per-step max-3-retries budget — it changes which model runs the next attempt, never whether an attempt is warranted.
- Re-entry after a reported PASS (the ONLY statement of this rule): a reported PASS does NOT close the work. If the user later says a step's result is wrong or its quality too low, re-enter that step's retry path under trigger (2), and that step's retry budget resets — the complaint opens a fresh cycle of up to 3 retries even if the earlier cycle was exhausted.
Cross-Provider Equivalence
When this skill runs outside the Anthropic model context, map the tier to the nearest model of the same class:
The mapping is by capability tier, not by name — exact names drift as vendors ship new models. Every rule above is expressed in tiers, so on another provider: map tier → your model of that class, then apply the selection, pairing and escalation rules unchanged.
Process
Setup: Create Reports Directory
Before starting, ensure the reports directory exists:
Report naming convention: .specs/reports/{task-name}-step-{N}-{YYYY-MM-DD}.md
Where:
{task-name}- Derived from task description (e.g.,user-dto-refactor){N}- Step number{YYYY-MM-DD}- Current date
Note: Implementation outputs go to their specified locations; only judge verification reports go to .specs/reports/
Phase 1: Task Analysis and Decomposition
Resolve configuration first: STRICT_MODE = --strict present || false. Strip all flags from the task text — never pass them into sub-agent prompts.
Analyze the task systematically using Zero-shot Chain-of-Thought reasoning:
Decomposition Guidelines:
Decomposition Output Format:
Phase 2: Model Selection for Each Subtask
Assess every subtask on the three axes below, then read its tier straight off the Selection Rules table — tiers are chosen per step, never once for the whole run.
- Scope — one file, one component, or multiple files?
- Complexity — mechanical edit, established pattern, or novel/intricate logic?
- Risk — isolated and reversible, internal, or critical per the exhaustive list in the Selection Rules
opusrow?
For each step, state the three findings, the chosen tier, and a one-line justification before dispatching it. Then apply Role Pairing — which governs in full, including its --model override — to decide that step's meta-judge tier.
Domain Expertise Check: "Does this subtask match a specialized agent profile?"
- Development: implementation, refactoring, bug fixes
- Architecture: system design, pattern selection
- Documentation: API docs, comments, README updates
- Testing: test generation, test updates
Specialized Agent: Specialized agent list depends on project and plugins that are loaded. Common agents from the sdd plugin include: sdd:developer, sdd:researcher, sdd:software-architect, sdd:tech-lead, sdd:business-analyst, sdd:code-explorer, sdd:code-reviewer, sdd:tech-writer. If the appropriate specialized agent is not available, fallback to a general agent without specialization.
Decision: Use specialized agent when subtask clearly benefits from domain expertise AND complexity justifies the overhead (not for haiku-tier steps).
Selection Output Format:
Phase 3: Sequential Execution with Parallel Meta-Judge and Judge Verification
Execute subtasks one by one. For each step, dispatch a meta-judge AND implementation agent in parallel, then verify with an independent judge using the meta-judge's specification. Iterate if needed, then pass context forward.
Execution Flow per Step:
3.1 Context Passing Protocol
After each subtask completes, extract relevant context for subsequent steps:
Context to pass forward:
- Files modified (paths only, not contents)
- Key changes made (summary)
- New interfaces/APIs introduced
- Decisions made that affect later steps
- Warnings or considerations for subsequent steps
Context filtering:
- Pass ONLY information relevant to remaining subtasks
- Do NOT pass implementation details that don't affect later steps
- Keep context summaries concise (max 200 words per step)
Context Size Guideline: If cumulative context exceeds ~500 words, summarize older steps more aggressively. Sub-agents can read files directly if they need details.
Example of Context Accumulation (Concrete):
3.2 Sub-Agent Prompt Construction
For each subtask, construct the prompt with these mandatory components:
3.2.1 Zero-shot Chain-of-Thought Prefix (REQUIRED - MUST BE FIRST)
3.2.2 Task Body
3.2.3 Self-Critique Suffix (REQUIRED - MUST BE LAST)
3.3 Parallel Meta-Judge Dispatch
CRITICAL: For each step, dispatch the meta-judge AND implementation agent in parallel in a single message with two Task tool calls. The meta-judge MUST be the first tool call in the message so it can observe artifacts before the implementation agent modifies them.
Both agents run as foreground agents. Wait for BOTH to complete before proceeding to judge dispatch.
Meta-Judge Prompt (per step):
Dispatch Example
Send BOTH Task tool calls in a single message. Meta-judge first, implementation second:
Wait for BOTH to return before proceeding to judge dispatch.
3.4 Judge Verification Protocol
After BOTH meta-judge and implementation agent complete, dispatch an independent judge to verify the step using the meta-judge evaluation specification.
CRITICAL: Provide to the judge EXACT meta-judge's evaluation specification YAML, do not skip or add anything, do not modify it in any way, do not shorten or summarize any text in it!
3.4.1 Analyze the Pre-existing Changes Section
Before dispatching the judge for each step, assess whether there are pre-existing changes in the codebase that the judge needs to be aware of. The "Pre-existing Changes" section prevents the judge from confusing prior modifications with the current step's implementation agent's work.
When to include:
- Previous steps' changes from the SAME do-in-steps run (steps 1..N-1 when judging step N) — this is the most common case in sequential execution. When running step N, the judge MUST know about changes from steps 1..N-1 as pre-existing. Each completed step's output (files created/modified, key changes) becomes pre-existing context for subsequent step judges.
- Previous do-in-steps or do-and-judge task runs completed earlier in the same session
- User's manual modifications made before invoking the skill (visible from conversation context or in git)
- Changes from other tools or agents that ran before this task
When to omit:
- This is step 1 with no known prior changes (no earlier session tasks, no user modifications) — omit the section entirely
- On retries within the SAME step, do NOT include the implementation agent's own previous attempt as "pre-existing changes" — those are part of the current step's iteration cycle
Content guidelines:
- Use a high-level summary: task description, list of affected files/modules, general nature of changes (created, modified, deleted)
- Do NOT include code blocks, diffs, or line-level details — keep it concise
- Label each source clearly: "Step 1: {description}", "Step 2: {description}", "User modifications (before current task)", etc.
- If multiple sources of pre-existing changes exist, use separate subsections for each (one per completed step, plus any external sources)
- Leverage the Context Passing Protocol output (section 3.1) — the "Completed Steps Summary" already tracks what each step produced
CRITICAL: avoid reading full codebase or git history, just use high-level git diff/status to determine which files were changed, or use conversation context and completed step summaries to determine pre-existing changes.
Prompt template for step judge:
Implementation Output
{Path to files modified by implementation agent} {Context for Next Steps section from implementation agent}
Instructions
Follow your full judge process as defined in your agent instructions!
Output
CRITICAL: You must reply with this exact structured evaluation report format in YAML at the START of your response!
Use Task tool:
- description: "Judge Step {N}/{total}: {subtask_name}"
- prompt: {judge verification prompt with exact meta-judge specification YAML, and Pre-existing Changes section if applicable}
- model: {judge model — the user's
--modelif one was passed; otherwise MUST equal this step's current implementation model, including after escalation} - subagent_type: "sadd:judge"
-
Dispatch meta-judge AND implementation agent IN PARALLEL (single message, 2 tool calls): Tool call 1 (meta-judge — MUST be first): Use Task tool: - description: "Meta-judge Step {N}/{total}: {subtask_name}" - prompt: {meta-judge prompt with step requirements and context} - model: {meta-judge model — the user's
--modelif one was passed; otherwise the same tier as this step's implementation, or one tier up per Role Pairing} - subagent_type: "sadd:meta-judge"Tool call 2 (implementation): Use Task tool: - description: "Step {N}/{total}: {subtask_name}" - prompt: {constructed prompt with CoT + task + previous context + self-critique} - model: {implementation model — the user's
--modelif one was passed; otherwise the model selected for this step} - subagent_type: "{selected agent type}" -
Wait for BOTH to complete. Collect outputs:
- From meta-judge: Extract evaluation specification YAML
- From implementation: Parse "Context for Next Steps" section, note files modified
-
Dispatch judge sub-agent (with this step's meta-judge specification): Use Task tool:
- description: "Judge Step {N}/{total}: {subtask_name}"
- prompt: {judge verification prompt with step requirements, implementation output, and meta-judge specification YAML}
- model: {judge model — the user's
--modelif one was passed; otherwise MUST equal this step's current implementation model, including after escalation} - subagent_type: "sadd:judge"
-
Parse judge verdict (DO NOT read full report): Extract from judge reply:
- VERDICT: PASS or FAIL
- SCORE: X.X/5.0
- ISSUES: List of problems (if any)
- IMPROVEMENTS: List of suggestions (if any)
-
Decision based on verdict:
If score ≥4.0: → VERDICT: PASS → Proceed to next step with accumulated context → Include IMPROVEMENTS in context as optional enhancements
If 3.0 ≤ score <4.0 and NOT STRICT_MODE: → Apply the Iteration Discretion Rule (3.6) → accepted → VERDICT: PASS (report outstanding issues and proceed) → declined → VERDICT: FAIL → go to "Check retry count" below
Otherwise (score <3.0, or score <4.0 with STRICT_MODE): → VERDICT: FAIL → Check retry count for this step
If retries < 3: → Decide this step's retry tier per "3.5.1 Model Escalation on Retry" below (per the Escalation Rule — bump BOTH, unless its sole hold exception applies) → Dispatch retry implementation agent at that tier with: - Original step requirements - Judge's ISSUES list as feedback - Path to judge report for details - Instruction to fix specific issues → Return to judge verification with SAME meta-judge specification from this step, dispatching the judge at the retry tier (judge always matches implementation) → Do NOT re-run meta-judge for retries, and do NOT re-tier it
If retries ≥ 3: → Escalate to user (see Error Handling) → Do NOT proceed to next step
-
Proceed to next subtask with accumulated context → Next step gets a NEW meta-judge dispatched in parallel with its implementation agent
3.6 Iteration Discretion Rule
Your main task is to COMPLETE the task within target quality, and iteration effort MUST stay proportionate to each step's size. Two failure modes are equally real:
- Burning retries and context on nitpicks so the overall task never completes → the task is failed.
- Accepting a step whose quality is genuinely too poor to be considered complete → an even worse failure.
Apply to every judge score:
score < 3.0→ FAIL, unconditionally. No discretion. Retry with judge feedback until the step passes or max retries is reached.3.0 <= score < 4.0→ discretion band. ONLY inside this band MAY you decide that a step below the4.0target is acceptable. The fixed4.0target puts the effective floor at3.0, so no separate bounded-drop guard is needed.- Inside the band, when the outstanding issues are ONLY low/medium priority (any High or Critical finding removes discretion entirely) AND none of them breaks a target requirement of the step or causes a meaningful defect (i.e. they are nitpicks), you MUST reason FIRST — before dispatching a retry — about whether another attempt is worth the time and context cost.
- At most ONE nitpick-driven retry, and it counts against the retry budget. If it again surfaces only nitpicks, you MUST mark the step PASS (
ACCEPTED), carry the outstanding issues forward in the accumulated context, report them in the final summary, and continue with the next step. If it returns a score below3.0, the unconditional-FAIL rule applies instead. - You MUST be critical, NOT lenient. Stopping short of target MUST be an intentional decision grounded in the absence of real, requirement-breaking issues — later steps build on this one. A genuine blocking issue that prevents completing the step within max retries MUST be escalated as a failure, never papered over.
- If
STRICT_MODEis true, this whole rule is DISABLED: stop only whenscore >= 4.0or max retries is reached.--strictchanges nothing else — the4.0target, the max-retry limit, the< 3.0unconditional FAIL and meta-judge/judge dispatch are unaffected.
Phase 4: Final Summary and Report
After all subtasks complete and pass verification, reply with a comprehensive report:
Error Handling
If Judge Verification Fails (Score <4.0)
The judge-verified iteration loop handles most failures automatically (a score in 3.0..4.0 is FAIL only after the Iteration Discretion Rule declines to accept it):
If Step Fails After Max Retries
When a step fails judge verification three times:
- STOP - Do not proceed with broken foundation
- Report - Provide failure analysis:
- Original step requirements
- All judge verdicts and scores
- Persistent issues across retries
- Escalate - Present options to user:
- Provide additional context/guidance for retry
- Re-run this step at the next model tier up (omit this option if the step already reached
opus) - Modify step requirements
- Skip step (if optional)
- Abort and report partial progress
- Wait - Do NOT proceed without user decision
Escalation Report Format:
Never:
- Continue past a failed step after max retries
- Skip judge verification to "save time"
- Ignore persistent issues across retries
- Make assumptions about what might have worked
If Context is Missing
- Do NOT guess what previous steps produced
- Re-examine previous step output for missing information
- Check judge reports - they may have noted missing elements
- Dispatch clarification sub-agent if needed to extract missing context
- Update context passing for future similar tasks
If Steps Conflict
- Stop execution at conflict point
- Analyze: Was decomposition incorrect? Are steps actually dependent?
- Check judge feedback - judges may have flagged integration issues
- Options:
- Re-order steps if dependency was missed
- Combine conflicting steps into one
- Add reconciliation step between conflicting steps
Examples
Example 1: Sequential Steps Building on Each Other (Pre-existing Changes from Previous Steps)
Input:
Phase 1 - Decomposition:
Phase 2 - Model Selection:
Phase 3 - Execution with Pre-existing Changes Accumulation:
Final Summary:
- Total Agents: 10 (3 meta-judges + 3 implementations + 0 retries + 3 judges)
- Pre-existing Changes Progression:
- Step 1 judge: None
- Step 2 judge: Step 1 output (2 files)
- Step 3 judge: Steps 1+2 output (5 files)
- All Judge Scores: 4.2, 4.4, 4.1
Example 2: User-Modified Codebase + Sequential Steps (Mixed Pre-existing Changes Sources)
Scenario:
The user has been working on a payment processing module during the conversation. They modified several files (added a new PaymentGateway interface, updated configuration) before invoking do-in-steps.
Input:
Phase 1 - Decomposition:
Phase 2 - Model Selection:
Both steps land on opus because the domain earns it, not because the run does — see Example 3 for a chain that mixes tiers.
Phase 3 - Execution with Mixed Pre-existing Changes:
Final Summary:
- Total Agents: 7 (2 meta-judges + 2 implementations + 0 retries + 2 judges)
- Pre-existing Changes Progression:
- Step 1 judge: User modifications (3 files)
- Step 2 judge: User modifications (3 files) + Step 1 output (2 files)
- All Judge Scores: 4.3, 4.5
Example 3: Multi-file Refactoring with Escalation
Input:
Phase 1 - Decomposition:
Phase 2 - Model Selection:
Phase 3 - Execution with Escalation (each step has parallel meta-judge + implementation):
Total Agents: 20 (5 meta-judges + 5 implementations + 5 retries + 5 judges)
Best Practices
Task Decomposition
- Be explicit: Each subtask should have a clear, verifiable outcome
- Define verification points: What should the judge check for each step?
- Minimize steps: Combine related work; don't over-decompose
- Validate dependencies: Ensure each step has what it needs from previous steps
- Plan context: Identify what context needs to pass between steps
Model Selection
The rules govern in the Model Selection Policy; these are the habits that make them stick:
- Justify out loud, per step - state scope, complexity and risk plus the resulting tier before dispatching each step; this is the highest-leverage decision in the run
opusis earned, never a hedge - resolve every overlap and tie by the Selection Rules precedence and tie-breaker, never by instinct- Tier each step on its own merits - a chain mixes tiers; a neighbouring step's tier (or an escalated one) is not evidence about this step
- One tier across roles - raise only the criteria-setter (meta-judge), and only for a non-obvious step (Role Pairing)
- Escalate on evidence - a clearly-too-low attempt or a user quality complaint (Escalation Rule), scoped to the failing step
Context Passing Guidelines
Keep context focused:
- Pass what the next step NEEDS to build on
- Omit internal details that don't affect subsequent steps
- Highlight patterns/conventions to maintain consistency
- Include judge IMPROVEMENTS as optional enhancements
- Track pre-existing changes - Pass context about prior modifications (including previous steps) to the judge to prevent attribution confusion
Meta-Judge + Judge Verification
- Never skip meta-judge - Tailored evaluation criteria produce better judgments than generic ones
- One meta-judge per step - Each step gets its own meta-judge dispatched in parallel with implementation
- Reuse meta-judge spec across retries within a step - On retry, reuse the same step's meta-judge specification; do NOT re-run meta-judge
- New meta-judge for each new step - Different steps have different requirements, so each gets a fresh meta-judge
- Meta-judge FIRST in parallel dispatch - Always the first tool call in the message
- Parse only headers from judge - Don't read full reports to avoid context pollution
- Include CLAUDE_PLUGIN_ROOT - Both meta-judge and judge need the resolved plugin root path
- Meta-judge YAML - Pass only the meta-judge YAML to the judge, do not add any additional text or comments to it!
- After self-critique: Judge reviews work that already passed internal verification
- Independent verification: Judge is different agent than implementer
- Structured output: Always parse VERDICT/SCORE from reply, not full report
- Max retries: 3 attempts before escalating to user
- Feedback loop: Pass judge ISSUES to retry implementation agent
- Return to judge verification with same step's meta-judge specification on retry
Quality Assurance
- Two-layer verification: Self-critique (internal) + Judge (external)
- Self-critique first: Implementation agents verify own work before submission
- External judge second: Independent judge catches blind spots self-critique misses
- Iteration loop: Retry with feedback until passing or max retries
- Proportionate iteration: Apply the Iteration Discretion Rule — at most ONE nitpick-driven retry, never below
3.0, disabled by--strict - Chain validation: Judges check integration with previous steps
- Escalation: Don't proceed past failed steps - get user input
- Final integration test: After all steps, verify the complete change works together
Context Format Reference
Implementation Agent Output Format
Judge Verdict Format (Structured Header)
Judge Verdict Format (FAIL Example)
Key Insight: Complex tasks with dependencies benefit from sequential execution where each step operates in a fresh context while receiving only the relevant outputs from previous steps. Per-step meta-judge evaluation specifications ensure tailored evaluation criteria specific to each step's requirements, while running in parallel with implementation for speed. External judge verification catches blind spots that self-critique misses, while the iteration loop (reusing the same step's meta-judge spec) ensures quality before proceeding. This prevents both context pollution and error propagation.

