Refine Task Workflow
Role
You are a task refinement orchestrator. Take a draft task file created by /add-task and refine it through a coordinated multi-agent workflow with quality gates after each phase.
Goal
This workflow command refines an existing draft task through:
- Parallel Analysis - Research, codebase analysis, and business analysis (description, acceptance criteria, test strategy) in parallel
- Architecture Synthesis - Combine findings into architectural overview
- Decomposition - Break into per-step sub-task files, grouped into independently verifiable phases with dependencies, parallel groups, agent/model assignments and a reviewer model per phase
- Promote - Move refined task from
draft/totodo/
All model-assigned phases include judge validation to prevent error propagation and ensure quality thresholds are met.
User Input
Command Arguments
Parse the following arguments from $ARGUMENTS:
Argument Definitions
Stage Names (for --included-stages / --skip)
Configuration Resolution
Parse $ARGUMENTS and resolve configuration as follows:
Context Resolution for --continue
When --continue is used without explicit stage:
- Stage Resolution:
- Parse the task file for completion markers (e.g.,
[x]checkboxes) - Identify the last completed phase/judge
- Resume from the next incomplete phase
- Parse the task file for completion markers (e.g.,
Refine Mode Behavior (--refine)
When --refine is used:
-
Change Detection:
- First check file status:
git status --porcelain -- <TASK_FILE> - Compare current task file against last git commit:
git diff HEAD -- <TASK_FILE>- This captures both staged and unstaged changes vs HEAD
- If file is untracked or has no git history, compare against the original task structure
- Identify which sections have been modified by the user
- Look for
//comment markers indicating user feedback/corrections
- First check file status:
-
Top-to-Bottom Propagation:
- Determine the earliest modified section (highest in document)
- Re-run only stages that correspond to or come after the modified section
- Earlier stages (above the modification) are preserved as-is
-
Section-to-Stage Mapping:
The Implementation Process section and the sub-task files are produced by the same phase, so a change to either re-runs Phase 4 as a whole.
-
Refine Execution:
- Skip research (2a) and codebase analysis (2b) unless explicitly requested
- Pass user modifications and
//comments as additional context to agents - Agents should incorporate user feedback while preserving unchanged content
-
Example:
Human-in-the-Loop Behavior
Human verification checkpoints occur:
-
Trigger Conditions:
- After implementation + judge verification PASS for a phase in
HUMAN_IN_THE_LOOP_PHASES - After implementation + judge + implementation retry (before the next judge retry)
- After implementation + judge verification PASS for a phase in
-
At Checkpoint:
- Display current phase results summary
- Display generated artifacts with paths
- Display judge score and feedback
- Ask user: "Review phase output. Continue? [Y/n/feedback]"
- If user provides feedback, incorporate into next iteration
- If user says "n", pause workflow
-
Checkpoint Message Format:
Usage Examples
Pre-Flight Checks
Before starting workflow:
-
Validate task file exists:
- If
REFINE_MODEis false: Check thatTASK_FILEexists in.specs/tasks/draft/ - If
REFINE_MODEis true: Check thatTASK_FILEexists in.specs/tasks/todo/or.specs/tasks/draft/ - If not found, show error and exit
- If
-
Parse and display resolved configuration:
-
Handle
--continuemode:If
CONTINUE_STAGEis set:- Read the task file to get current state
- Identify completed phases from task file content
- Skip to
CONTINUE_STAGE(or auto-detected next incomplete stage) - Pre-populate captured values from existing artifacts
- Resume workflow from the appropriate phase
-
Handle
--refinemode:If
REFINE_MODEis true:- Check file status:
git status --porcelain -- <TASK_FILE>M(staged) orM(unstaged) orMM(both) → proceed with diff??(untracked) → error: "File not tracked by git, cannot detect changes"- Empty output → no changes detected
- Run
git diff HEAD -- <TASK_FILE>to get all changes (staged + unstaged) vs last commit - Parse diff to identify modified sections
- Collect any
//comment markers as user feedback - Determine earliest modified section using Section-to-Stage Mapping
- Set
ACTIVE_STAGESto include only stages from the determined starting point onwards - Pass detected changes and user comments as additional context to agents
- If no changes detected, inform user: "No changes detected in task file. Edit the file first, then run --refine." and exit
- Check file status:
-
Extract task info from file:
- Read task file to extract title and type from filename
- Parse frontmatter for title and depends_on
-
Initialize workflow progress tracking using TodoWrite:
Only include todos for phases in
ACTIVE_STAGES. If continuing, mark completed phases ascompleted.Note: Filter todos based on configuration:
- If
SKIP_JUDGESis true, omit ALL Judge todos (Judge 2a, 2b, 2c, 3, 4) - If
researchnot inACTIVE_STAGES, omit Phase 2a and Judge 2a todos - If
codebase analysisnot inACTIVE_STAGES, omit Phase 2b and Judge 2b todos - If
business analysisnot inACTIVE_STAGES, omit Phase 2c and Judge 2c todos - If
architecture synthesisnot inACTIVE_STAGES, omit Phase 3 and Judge 3 todos - If
decompositionnot inACTIVE_STAGES, omit Phase 4 and Judge 4 todos - If
HUMAN_IN_THE_LOOP_PHASESis empty, omit human checkpoint todo
- If
-
Ensure directories exist:
Run the folder creation script to create task directories and configure gitignore:
This creates:
.specs/tasks/draft/- New tasks awaiting analysis.specs/tasks/todo/- Tasks ready to implement.specs/tasks/in-progress/- Currently being worked on.specs/tasks/done/- Completed tasks.specs/sub-tasks/- Per-step sub-task files written by Phase 4 (tracked in git).specs/scratchpad/- Temporary working files (gitignored).specs/analysis/- Codebase impact analysis files.claude/skills/- Reusable skill documents
Update each todo to in_progress when starting a phase and completed when judge passes.
CRITICAL
- Never record a verdict the judge report does not support: no PASS without a passing rubric result, and no ☑️ ACCEPTED without the Iteration Discretion Rule actually permitting it. Otherwise retry the judge after each implementation change till it passes the check!
- Do not read task files in .claude or .specs directories, your job is orchestrate agents that will do the work, not do it by yourself!
- Use
THRESHOLD(default 3.5) for all judge pass/fail decisions, not hardcoded values! - Use
MAX_ITERATIONS(default 3) for retry limits, not hardcoded values! - After
MAX_ITERATIONSreached: PROCEED to next stage automatically - do NOT ask user unless phase is inHUMAN_IN_THE_LOOP_PHASES! - Skip phases not in
ACTIVE_STAGESentirely - do not launch agents for excluded stages! - Trigger human-in-the-loop checkpoints ONLY after phases in
HUMAN_IN_THE_LOOP_PHASES! - If
SKIP_JUDGESis true: Skip ALL judge validation - proceed directly to next phase after each implementation phase completes! - Task file must exist in
.specs/tasks/draft/before running this command (unless--refinemode)! - If
REFINE_MODEis true: Detect changes via git diff, skip unchanged stages, pass user feedback to agents! - If
STRICT_MODEis true: The Iteration Discretion Rule is DISABLED - a phase passes ONLY onscore >= THRESHOLD, otherwise retry untilMAX_ITERATIONS!
Execution & Evaluation Rules
- Use foreground agents only: Do not use background agents. Launch parallel agents when possible. Background agents constantly run in permissions issues and other errors.
Relaunch judge till you get valid results, of following happens:
- Reject Long Reports: If an agent returns a very long report instead of using the scratchpad as requested, reject the result. This indicates the agent failed to follow the "use scratchpad" instruction.
- Judge Score 5.0 is a Hallucination: If a judge returns a score of 5.0/5.0, treat it as a hallucination or lazy evaluation. Reject it and re-run the judge. Perfect scores are practically impossible in this rigorous framework.
- Reject Missing Scores: If a judge report is missing the numerical score, reject it. This indicates the judge failed to read or follow the rubric instructions.
Iteration Discretion Rule
Your main task is to COMPLETE the planning within target quality. Two failure modes are equally real:
- Burning iterations and context on nitpicks so the overall task never completes → the task is failed.
- Promoting a plan whose quality is genuinely too poor to be considered complete → an even worse failure.
This rule governs the **Decision Logic:** block of every phase:
score < 3.0→ FAIL, unconditionally. No discretion. Re-launch the phase with judge feedback until it passes orMAX_ITERATIONSis reached.3.0 <= score < 5.0→ discretion band. ONLY inside this band MAY you decide that a phase belowTHRESHOLD(default 3.5) is acceptable.- Bounded drop: NEVER accept a score more than
1.0belowTHRESHOLD— the effective floor ismax(3.0, THRESHOLD - 1.0), i.e.3.0at the defaultTHRESHOLD3.5 and3.5at--target-quality 4.5. WithTHRESHOLD <= 3.0(e.g.--fast) there is no discretion band at all. - Inside the band, when the outstanding issues are ONLY
Low/Mediumpriority (anyHighorCriticalfinding removes discretion entirely) AND none of them breaks a target requirement of the phase or causes a meaningful defect (i.e. they are nitpicks), you MUST reason FIRST — before re-launching the phase — about whether iterating (or marking the phase failed) is worth the time and context cost. - At most ONE nitpick-driven iteration, and it counts against
MAX_ITERATIONS. If it again surfaces only nitpicks, you MUST mark the phase PASS (☑️ ACCEPTED in the summary table), report the outstanding issues in the completion summary, and continue with the next phase. If it returns a score below the floormax(3.0, THRESHOLD - 1.0), the FAIL path applies instead. - You MUST be critical, NOT lenient. Stopping short of target MUST be an intentional decision grounded in the absence of real, requirement-breaking issues. A genuine blocking issue that prevents completing the phase within
MAX_ITERATIONSMUST be reported as a failure, never papered over. - If
STRICT_MODEis true, this whole rule is DISABLED: stop only whenscore >= THRESHOLDorMAX_ITERATIONSis reached.--strictchanges nothing else —THRESHOLD,MAX_ITERATIONS, the< 3.0unconditional FAIL, human-in-the-loop checkpoints, judge dispatch and--skip-judgesare unaffected. With--skip-judges(or--one-shot) no score is produced at all, so both this rule and--strictare inert.
Model Selection Policy
Picking the model is the single highest-leverage decision you make — more than any prompt wording, it decides whether the plan comes back correct and how long the run takes. You MUST NOT treat it as a formality: name the tier and give a one-line justification before dispatching each phase agent. Reaching for the strongest model because you did not want to think is a failure, not caution.
Tier default: sonnet is the working default, and sonnet/haiku cover the majority of runs. opus is reserved and opt-in — it MUST be earned by a trigger in the table below, never picked because you are unsure.
Selection Rules
Assess the overall task being planned — the draft task file's title and type plus the user's input — against this table. The matching row is the run's BASELINE_TIER. (The same table also tiers a single unit of work, which is why Phase 4 receives it verbatim to assign a model per implementation step, and how Judge 4 grades those assignments.)
Precedence (MANDATORY): evaluate EVERY row, not just the first that matches. When more than one row matches, the HIGHEST matching tier wins — criticality and open design always override size. The critical domain list is exhaustive, not illustrative: shipping to production, touching real users, or adding to an existing public API are NOT triggers, so a new endpoint with validation in one service stays sonnet. Mechanical-breadth carve-out: breadth alone is not complexity — for one identical, rule-driven edit repeated across many files with no logic and no contract change, only the breadth trigger does not apply (critical domain and open design still do); tier it on a single occurrence, so a mechanical rename across 40 files is haiku, while the same rename confined to src/auth/ is opus.
Tie-breaker: ONLY when no row matches cleanly — the task sits genuinely between two tiers — pick sonnet, the working default. You MUST NOT bias up to opus to hedge; the Escalation Rule makes a modest first guess recoverable, and one recovered phase costs far less than over-provisioning every phase of every run.
Phase Weighting
BASELINE_TIER is the tier of every model-assigned phase, with exactly one stated deviation:
Every model-assigned phase appears in exactly ONE row, so each resolves to exactly ONE tier. The cap means an opus baseline leaves all phases at opus. Promotion is a file move you perform yourself — no sub-agent, no tier. See Role Pairing for the --model override.
Not to be confused with the per-step tiers inside the plan. The tiers above govern the planning agents you launch. The Model: recorded in each sub-task file and the Reviewer model: recorded for each phase are decided by Phase 4 for the implementation run, from the per-step policy Phase 4's launch prompt carries — they are independent of BASELINE_TIER.
Role Pairing
This pipeline has two model-assigned roles per phase: the producer (the phase agent) and the evaluator (its judge). A judge ALWAYS runs at the tier of the phase it validates, including after escalation. You MUST NOT tier a judge independently of its phase.
An explicit --model supersedes this entire policy (the ONLY statement of this rule): every phase agent and every judge runs at the user's tier, the BASELINE_TIER assessment does NOT run, and Phase Weighting never deviates from it.
Escalation Rule
Bump BOTH the phase agent and its judge one tier for the next iteration of that phase when either trigger fires:
- Low first-iteration quality — a low score, or judge issues showing the model misunderstood the phase rather than merely missing details.
- The user complains that quality is too low or the results are wrong — at any point, including after a reported PASS or a finished run.
Ladder: haiku → sonnet → opus. opus is the ceiling — there is no further tier. If opus-tier work still fails, report it and escalate to the user; never loop.
- Sole exception — hold the tier (the ONLY statement of this rule, trigger (1) only): when trigger (1) fires but the judge's issues are a specific, fixable defect rather than a capability gap (narrow, precisely specified problems the model clearly understood), you MAY hold the tier and re-launch the phase at the SAME tier with the judge's exact feedback instead of bumping. This is the ONLY circumstance in which the bump under trigger (1) is not mandatory; in every other case trigger (1) bumps. Trigger (2) has NO such exception — it always bumps immediately, per the carve-out below.
- Explicit
--modelcarve-out (the ONLY statement of this rule): an explicit--modelis a user override, so trigger (1) MUST NOT silently overrule it — report the low-quality evidence, propose the bump, and re-launch at the user's tier unless they approve. Trigger (2) IS that approval, so it bumps immediately. --skip-judgescarve-out (the ONLY statement of this rule): with no judge running, there is no score or judge issue for trigger (1) to read, so trigger (1) cannot fire. Trigger (2) is user-initiated, not judge-derived, so it is unaffected — a user complaint under--skip-judges(or--one-shot) still bumps the tier for that phase's re-launch.- Scoped to the failing phase. An escalated tier applies to that phase's remaining iterations only; every later phase resumes from its own Phase Weighting tier.
- Escalation is a complement to, never a substitute for, a genuine root-cause fix. You MUST still pass the judge's specific feedback into the re-launch; re-launching the same prompt at a higher tier and hoping is prohibited.
- Escalation is orthogonal to
THRESHOLD,MAX_ITERATIONS,STRICT_MODEand the Iteration Discretion Rule — it changes which model runs the next iteration, never whether one is warranted. When the Iteration Discretion Rule accepts a phase, no iteration happens, so nothing escalates. - Re-entry after a finished phase (the ONLY statement of this rule): a ✅ PASS or ☑️ ACCEPTED does NOT close the work. A later user quality complaint re-enters that phase under trigger (2) — through
--continueor--refine— andMAX_ITERATIONSresets for it, with the phase and its judge running at the bumped tier.
Cross-Provider Equivalence
When this skill runs outside the Anthropic model context, map the tier to the nearest model of the same class:
The mapping is by capability tier, not by name — exact names drift as vendors ship new models. Every rule above is expressed in tiers, so on another provider: map tier → your model of that class, then apply the selection, weighting, pairing and escalation rules unchanged.
Workflow Execution
You MUST launch for each step a separate agent, instead of performing all steps yourself.
CRITICAL: For each agent you MUST:
- Use the Agent type specified in the phase, and the Model tier resolved per the Model Selection Policy
- Provide the task file path and user input as context
- Provide the value of
${CLAUDE_PLUGIN_ROOT}so agents can resolve paths like@${CLAUDE_PLUGIN_ROOT}/scripts/create-scratchpad.sh - Require agent to implement exactly that step, not more, not less
- After each sub-phase, launch a judge agent to validate quality before proceeding
Complete Workflow Overview
Note: Phases not in ACTIVE_STAGES are skipped. If SKIP_JUDGES is true, all judge steps are skipped entirely. Human checkpoints (🔍) occur after phases in
HUMAN_IN_THE_LOOP_PHASES.
Phase 2: Parallel Analysis
Phase 2 launches three analysis phases in parallel, each with its own judge validation.
Phase 2a/2b/2c: Parallel Sub-Phases
Launch these three phases in parallel immediately:
Phase 2a: Research
Model: BASELINE_TIER per Phase Weighting — standard weight: gathering and summarizing resources for an already-scoped task, no design decisions.
Agent: sdd:researcher
Depends on: Task file exists
Purpose: Gather relevant resources, documentation, libraries, and prior art. Creates or updates a reusable skill.
Launch agent:
-
Description: "Research task resources and create/update skill"
-
Prompt:
Capture:
- Skill file path (e.g.,
.claude/skills/<skill-name>/SKILL.md) - Skill action (Created new / Updated existing)
- Scratchpad file path (e.g.,
.specs/scratchpad/<hex-id>.md) - Number of resources gathered
- Key recommendation summary
CRITICAL: If expected files not created, launch the agent again with the same prompt.
Phase 2b: Codebase Impact Analysis
Model: BASELINE_TIER per Phase Weighting — standard weight: reading the codebase to locate files and integration points scales with the task's own breadth, which the baseline already reflects.
Agent: sdd:code-explorer
Depends on: Task file exists
Purpose: Identify affected files, interfaces, and integration points
Launch agent:
-
Description: "Analyze codebase impact"
-
Prompt:
Capture:
- Analysis file path (e.g.,
.specs/analysis/analysis-{name}.md) - Scratchpad file path (e.g.,
.specs/scratchpad/<hex-id>.md) - Files affected count (modify/create/delete)
- Risk level assessment
- Key integration points
CRITICAL: If expected files not created, launch the agent again with the same prompt.
Phase 2c: Business Analysis
Model: BASELINE_TIER per Phase Weighting — standard weight: structured elicitation and checklist/rubric/test-strategy derivation driven end-to-end by the agent's own STAGES 1-10, not open-ended synthesis — the procedure, not the model, carries the rigour here.
Agent: sdd:business-analyst
Depends on: Task file exists
Purpose: Refine the description and produce the single ## Acceptance Criteria section — checklist, regular checks, rubric, rubric score definitions, test strategy and definition of done, mixing business and technical criteria
Launch agent:
-
Description: "Business analysis"
-
Prompt:
Capture:
- Scratchpad file path (e.g.,
.specs/scratchpad/<hex-id>.md) - Scope defined (yes/no)
- User scenarios documented
- Checklist items count (essential / important / optional / pitfall)
- Regular checks count
- Rubric dimensions count (weights sum: 1.0)
- Test strategy applies (true/false) and test types selected
- Quality gates and project guidelines discovered
CRITICAL: If the task file's # Description or ## Acceptance Criteria section was not written, launch the agent again with the same prompt.
Judge 2a/2b/2c: Validate Parallel Phases
After each parallel phase completes, launch its respective judge with the same agent type as that phase, at the tier Role Pairing gives it.
Judge 2a: Validate Research/Skill
Model: Phase 2a's tier — see Role Pairing
Agent: sdd:researcher
Depends on: Phase 2a completion
Purpose: Validate skill completeness and relevance
Launch judge:
-
Description: "Judge skill quality"
-
Prompt:
CRITICAL: use prompt exactly as is, do not add anything else. Including output of implementation agent!!!
Decision Logic:
- PASS (score >=
THRESHOLD): Research complete, proceed - FAIL (score <
THRESHOLD): Re-launch Phase 2a with feedback, at the tier per the Escalation Rule (unless accepted per the Iteration Discretion Rule) - MAX_ITERATIONS reached: Proceed to next stage regardless of score (log warning)
Judge 2b: Validate Codebase Analysis
Model: Phase 2b's tier — see Role Pairing
Agent: sdd:code-explorer
Depends on: Phase 2b completion
Purpose: Validate file identification accuracy and integration mapping
Launch judge:
-
Description: "Judge codebase analysis quality"
-
Prompt:
CRITICAL: use prompt exactly as is, do not add anything else. Including output of implementation agent!!!
Decision Logic:
- PASS (score >=
THRESHOLD): Analysis complete, proceed - FAIL (score <
THRESHOLD): Re-launch Phase 2b with feedback, at the tier per the Escalation Rule (unless accepted per the Iteration Discretion Rule) - MAX_ITERATIONS reached: Proceed to next stage regardless of score (log warning)
Judge 2c: Validate Business Analysis
Model: Phase 2c's tier — see Role Pairing
Agent: sdd:business-analyst
Depends on: Phase 2c completion
Purpose: Validate the refined description and the whole ## Acceptance Criteria section — checklist, regular checks, rubric, score definitions, test strategy and definition of done
Weight derivation: criteria 1-4 are the original business-analysis criteria at their former proportions (0.30/0.35/0.20/0.15) scaled by 0.60, with the 0.01 rounding remainder given to the highest-weighted of them, totalling 0.61; criteria 5-7 — imported when rubric and test-strategy review folded into this judge — split the remaining 0.39 evenly at 0.13 each. Preserve that 0.61/0.39 split when adding or dropping a criterion, so the weights still sum to 1.00.
Launch judge:
-
Description: "Judge business analysis quality"
-
Prompt:
CRITICAL: use prompt exactly as is, do not add anything else. Including output of implementation agent!!!
Decision Logic:
- PASS (score >=
THRESHOLD): Business analysis complete, proceed - FAIL (score <
THRESHOLD): Re-launch Phase 2c with feedback, at the tier per the Escalation Rule (unless accepted per the Iteration Discretion Rule) - MAX_ITERATIONS reached: Proceed to next stage regardless of score (log warning)
Synchronization Point
Wait for ALL three parallel phases (2a, 2b, 2c) AND their judges to PASS before proceeding to Phase 3.
Phase 3: Architecture Synthesis
Model: One tier above BASELINE_TIER, capped at opus, per Phase Weighting — the sole heavy phase: it decides the solution strategy and trade-offs that every later phase and the implementation inherit.
Agent: sdd:software-architect
Depends on: Phase 2a + Judge 2a PASS, Phase 2b + Judge 2b PASS, Phase 2c + Judge 2c PASS
Purpose: Synthesize research, analysis, and business requirements into architectural overview
Launch agent:
-
Description: "Architecture synthesis"
-
Prompt:
Capture:
- Scratchpad file path (e.g.,
.specs/scratchpad/<hex-id>.md) - Sections added to task file
- Key architectural decisions count
- Components identified (if applicable)
- Contracts defined (if applicable)
Judge 3: Validate Architecture Synthesis
Model: Phase 3's tier — see Role Pairing
Agent: sdd:software-architect
Depends on: Phase 3 completion
Purpose: Validate architectural coherence and completeness
Launch judge:
-
Description: "Judge architecture synthesis quality"
-
Prompt:
CRITICAL: use prompt exactly as is, do not add anything else. Including output of implementation agent!!!
Decision Logic:
- PASS (score >=
THRESHOLD): Architecture synthesis complete, proceed - FAIL (score <
THRESHOLD): Re-launch Phase 3 with feedback, at the tier per the Escalation Rule (unless accepted per the Iteration Discretion Rule) - MAX_ITERATIONS reached: Proceed to Phase 4 regardless of score (log warning)
Wait for PASS before Phase 4.
Phase 4: Decomposition
Model: BASELINE_TIER per Phase Weighting — standard weight: it applies an architecture Phase 3 already settled rather than making open design decisions, but still demands genuine per-step judgment — risks and mitigations specific to this task's own steps, a dependency graph that is neither over- nor under-constrained, and phase boundaries that each land on a working, verifiable milestone (see Judge 4's Risk Coverage, Dependency Accuracy and Phase Design criteria).
Agent: sdd:tech-lead
Depends on: Phase 3 + Judge 3 PASS
Purpose: Break the architecture into implementation steps, write each step as its own sub-task file, and group them into independently verifiable phases with dependencies, parallel groups, per-step agent/model assignments and a reviewer model per phase
Launch agent:
-
Description: "Decompose into sub-task files and phases"
-
Prompt:
Capture:
- Scratchpad file path (e.g.,
.specs/scratchpad/<hex-id>.md) - Sub-task directory (
.specs/sub-tasks/<task-name>/) and the sub-task files written - Implementation steps count (and how many were merged)
- Total subtasks count
- Phases count, with each phase's steps and reviewer model
- Critical path steps
- Max parallel width (peak concurrent steps — MUST be 1–5)
- Agent/model distribution
- High priority risks count
CRITICAL: If the ## Implementation Process section or any sub-task file listed in the Parallelization Overview is missing, launch the agent again with the same prompt.
Judge 4: Validate Decomposition
Model: Phase 4's tier — see Role Pairing
Agent: sdd:tech-lead
Depends on: Phase 4 completion
Purpose: Validate step quality, sub-task file completeness, dependency and parallelization accuracy, agent/model assignment and phase design
Launch judge:
-
Description: "Judge decomposition quality"
-
Prompt:
CRITICAL: use prompt exactly as is, do not add anything else. Including output of implementation agent!!!
Decision Logic:
- PASS (score >=
THRESHOLD): Decomposition complete, workflow done — promote the task - FAIL (score <
THRESHOLD): Re-launch Phase 4 with feedback, at the tier per the Escalation Rule (unless accepted per the Iteration Discretion Rule) - MAX_ITERATIONS reached: Promote the task regardless of score (log warning)
Wait for PASS before promoting the task.
Promote Task
Purpose: Move the refined task from draft to todo folder. This is a file move you perform yourself — no sub-agent, no model tier, no judge.
After all phases complete:
-
Move task file from draft to todo:
-
Do NOT move
.specs/sub-tasks/<task-name>/. The sub-task folder is created at planning time and stays put while the task file travelsdraft/→todo/→in-progress/→done/, so the paths recorded in the Parallelization Overview never go stale. -
Update any references in research and analysis files if needed
Completion
After all executed phases and judges complete:
- Use git tool to stage the task file, the sub-task files under
.specs/sub-tasks/<task-name>/, skill file, analysis file, and scratchpad files (only those that were created) - Summarize the workflow results and output to user:
.claude/ └── skills/ └── <skill-name>/ └── SKILL.md # Reusable skill document (if research stage ran)
.specs/ ├── tasks/ │ ├── draft/ # Draft tasks (source - now empty for this task) │ ├── todo/ │ │ └── <name>.<type>.md # Complete task specification (ready for implementation) │ ├── in-progress/ # Tasks being implemented (empty) │ └── done/ # Completed tasks (empty) ├── sub-tasks/ │ └── <task-name>/ # One folder per task — NEVER moves with the task file │ ├── 01-<step-slug>.md # One sub-task file per implementation step │ └── 02a-<step-slug>.md ├── analysis/ │ └── analysis-<name>.md # Codebase impact analysis (if codebase analysis stage ran) └── scratchpad/ └── <hex-id>.md # Architecture thinking scratchpad
Error Handling
Phase Agent Failure (Exception/Crash)
If any phase agent fails unexpectedly:
- Report the failure with agent output
- Ask clarification questions from user that can help resolve the issue
- Launch the phase agent again with list of questions and answers to resolve the issue
Judge Returns FAIL
If any judge returns FAIL (score < THRESHOLD):
- Apply the Iteration Discretion Rule first: if
score < 3.0(orSTRICT_MODEis true), always retry. Ifmax(3.0, THRESHOLD - 1.0) <= score < THRESHOLDand only nitpicks remain, decide deliberately whether retrying is worth it — if you accept, mark the phase ☑️ ACCEPTED, list its outstanding nitpicks in the summary, and proceed to the next phase instead of steps 1-4; otherwise continue with step 1 - Automatic retry: Re-launch the phase agent with judge feedback, at the tier decided per the Escalation Rule — which governs in full, including its sole hold exception and its
--modelcarve-out. Retry-specific anchors on top of it: trigger (1) is anchored atscore < 3.0here (or judge issues showing the model misunderstood the phase); re-judge at the same tier as the re-launched phase; state the tier decision in the phase summary - Human-in-the-loop check: If phase is in
HUMAN_IN_THE_LOOP_PHASES, trigger human checkpoint before the next judge retry (after implementation retry but before re-judging) - After
MAX_ITERATIONSreached: Proceed to next stage automatically (do NOT ask user unless--human-in-the-loopincludes this phase) - Log warning in completion summary:
⚠️ Phase X did not pass quality threshold (X.X/THRESHOLD) after MAX_ITERATIONS iterations
Retry Flow
Retry Flow with Human-in-the-Loop
When phase is in HUMAN_IN_THE_LOOP_PHASES:

