tree-of-thoughts
<task>
Execute complex reasoning tasks through systematic exploration of solution space, pruning unpromising branches, expanding viable approaches, and synthesizing the best solution.
</task>
<context>
This command implements the Tree of Thoughts (ToT) pattern for tasks requiring exploration of multiple solution paths before committing to full implementation. It combines creative sampling, meta-judge-generated evaluation specifications, multi-perspective evaluation, adaptive strategy selection, and evidence-based synthesis to produce superior outcomes.
Key benefits:
- Systematic exploration - Multiple agents explore different regions of the solution space
- Structured evaluation - Meta-judges produce tailored rubrics and criteria before judging
- Independent verification - Judges apply meta-judge specifications mechanically, reducing bias
- Adaptive strategy - Clear winners get polished, split decisions get synthesized, failures get redesigned
</context>
Pattern: Tree of Thoughts (ToT)
This command implements an eight-phase systematic reasoning pattern with meta-judge evaluation and adaptive strategy selection:
Process
Setup: Create Directory Structure
Before starting, ensure the directory structure exists:
Naming conventions:
- Proposals:
.specs/research/{solution-name}-{YYYY-MM-DD}.proposals.[a|b|c].md - Pruning:
.specs/research/{solution-name}-{YYYY-MM-DD}.pruning.[1|2|3].md - Selection:
.specs/research/{solution-name}-{YYYY-MM-DD}.selection.md - Evaluation:
.specs/reports/{solution-name}-{YYYY-MM-DD}.[1|2|3].md
Where:
{solution-name}- Derived from output path (e.g.,users-apifrom outputspecs/api/users.md){YYYY-MM-DD}- Current date
Note: Solutions remain in their specified output locations; only research and evaluation files go to .specs/
Phase 1: Exploration (Propose Approaches)
Launch 3 independent agents in parallel (recommended: Sonnet for speed):
- Each agent receives identical task description and context
- Each agent generates 6 high-level approaches (not full implementations)
- For each approach, agent provides:
- Approach description (2-3 paragraphs)
- Key design decisions and trade-offs
- Probability estimate (0.0-1.0)
- Estimated complexity (low/medium/high)
- Potential risks and failure modes
- Proposals saved to
.specs/research/{solution-name}-{date}.proposals.[a|b|c].md
Key principle: Systematic exploration through probabilistic sampling from the full distribution of possible approaches.
Prompt template for explorers:
Phase 1.5: Dispatch Pruning Meta-Judge
CRITICAL: Launch the pruning meta-judge in parallel with Phase 1 exploration agents. The meta-judge does not need exploration output to generate pruning criteria — it only needs the original task description.
The pruning meta-judge generates an evaluation specification (rubrics, checklist, scoring criteria) tailored to evaluating high-level proposals for pruning.
Prompt template for pruning meta-judge:
Dispatch:
Phase 2: Pruning (Vote for Top 3 Candidates)
Wait for BOTH Phase 1 exploration agents AND Phase 1.5 pruning meta-judge to complete before proceeding.
Launch 3 independent judges in parallel (recommended: Opus for rigor):
- Each judge receives ALL proposal files (from
.specs/research/) and the pruning meta-judge evaluation specification YAML - Judges evaluate each proposal against the meta-judge-generated pruning criteria
- Each judge produces:
- Scores for each proposal (with evidence)
- Vote for top 3 proposals to expand
- Rationale for selections
- Votes saved to
.specs/research/{solution-name}-{date}.pruning.[1|2|3].md
Key principle: Independent evaluation with meta-judge-generated criteria ensures consistent, tailored assessment without hardcoded weights.
CRITICAL: Provide to each judge the EXACT pruning meta-judge's evaluation specification YAML. Do not skip, add, modify, shorten, or summarize any text in it!
Prompt template for pruning judges:
Output
{.specs/research/{solution-name}-{date}.pruning.[1|2|3].md}
Instructions
Follow your full judge process as defined in your agent instructions!
CRITICAL: You must reply with this exact structured evaluation report format in YAML at the START of your response!
Use Task tool:
- description: "Pruning Judge {1|2|3}: {brief task summary}"
- prompt: {pruning judge prompt with exact meta-judge specification YAML}
- model: opus
- subagent_type: "sadd:judge"
Phase 3.5: Dispatch Evaluation Meta-Judge
CRITICAL: Launch the evaluation meta-judge in parallel with Phase 3 expansion agents. The meta-judge does not need expansion output to generate evaluation criteria — it only needs the original task description.
The evaluation meta-judge generates an evaluation specification (rubrics, checklist, scoring criteria) tailored to evaluating full solution implementations.
Prompt template for evaluation meta-judge:
Dispatch:
Phase 4: Evaluation (Judge Full Solutions)
Wait for BOTH Phase 3 expansion agents AND Phase 3.5 evaluation meta-judge to complete before proceeding.
Launch 3 independent judges in parallel (recommended: Opus for rigor):
- Each judge receives ALL solution files (solution.a.md, solution.b.md, solution.c.md) and the evaluation meta-judge specification YAML
- Judges evaluate against the meta-judge-generated evaluation criteria
- Each judge produces:
- Comparative analysis (which solution excels where)
- Evidence-based ratings (with specific quotes/examples)
- Final vote (which solution they prefer and why)
- Reports saved to
.specs/reports/{solution-name}-{date}.[1|2|3].md
Key principle: Multiple independent evaluations with meta-judge-generated specifications and explicit evidence reduce bias and catch different quality aspects.
CRITICAL: Provide to each judge the EXACT evaluation meta-judge's evaluation specification YAML. Do not skip, add, modify, shorten, or summarize any text in it!
CRITICAL: NEVER provide score threshold to judges. Judge MUST not know what threshold for score is, in order to not be biased!!!
Prompt template for evaluation judges:
Output
Write full report to: .specs/reports/{solution-name}-{date}.[1|2|3].md
CRITICAL: You must reply with this exact structured header format:
VOTE: [Solution A/B/C] SCORES: Solution A: [X.X]/5.0 Solution B: [X.X]/5.0 Solution C: [X.X]/5.0 CRITERIA:
- {criterion_1}: [X.X]/5.0
- {criterion_2}: [X.X]/5.0 ...
[Summary of your evaluation]
Instructions
Follow your full judge process as defined in your agent instructions!
CRITICAL: You must reply with this exact structured evaluation report format in YAML at the START of your response!
Use Task tool:
- description: "Evaluation Judge {1|2|3}: {brief task summary}"
- prompt: {evaluation judge prompt with exact meta-judge specification YAML}
- model: opus
- subagent_type: "sadd:judge"
Strategy 2: REDESIGN
When: All solutions scored <3.0/5.0 (fundamental issues across the board)
Process:
- Launch new agent to analyze the failure modes and lessons learned
- Return to Phase 3 (Expansion), provide to new implementation agents the lessons learned and new constraints
Note: If redesign fails twice, escalate to user for guidance.
Prompt template for new implementation:
Strategy 3: FULL_SYNTHESIS (Default)
When: No clear winner AND solutions have merit (scores >=3.0)
Process: Proceed to Phase 5 (Evidence-Based Synthesis)
Phase 5: Synthesis (Evidence-Based Combination)
Only executed when Strategy 3 (FULL_SYNTHESIS) selected in Phase 4.5
Launch 1 synthesis agent (recommended: Opus for quality):
- Agent receives:
- All solutions (from specified output location)
- All evaluation reports (from
.specs/reports/) - Selection rationale from pruning phase (from
.specs/research/)
- Agent analyzes:
- Consensus strengths (what multiple judges praised)
- Consensus weaknesses (what multiple judges criticized)
- Complementary elements where solutions took different approaches
- Agent produces final solution by:
- Copying superior sections when one solution clearly wins
- Combining approaches when hybrid is better
- Fixing identified issues that judges caught
- Documenting decisions (what was taken from where and why)
Key principle: Evidence-based synthesis leverages collective intelligence from exploration and evaluation.
Prompt template for synthesizer:
<output>
The command produces different outputs depending on the adaptive strategy selected:
Outputs (All Strategies)
-
Research directory:
.specs/research/(created if not exists)- Proposals:
.specs/research/{solution-name}-{date}.proposals.[a|b|c].md- High-level approaches with probabilities - Pruning:
.specs/research/{solution-name}-{date}.pruning.[1|2|3].md- Judge evaluations and votes - Selection:
.specs/research/{solution-name}-{date}.selection.md- Vote tallies and selected proposals
- Proposals:
-
Expansion outputs:
solution.a.md,solution.b.md,solution.c.md- Full implementations (in specified output location)
-
Reports directory:
.specs/reports/(created if not exists)- Evaluation:
.specs/reports/{solution-name}-{date}.[1|2|3].md- Final judge reports
- Evaluation:
-
Resulting solution:
{output_path}
Strategy-Specific Outputs
- SELECT_AND_POLISH: Polished solution based on winning solution, with targeted improvements
- REDESIGN: Do not stop; return to Phase 3 with lessons learned; eventually finishes at SELECT_AND_POLISH or FULL_SYNTHESIS
- FULL_SYNTHESIS: Synthesized solution combining best elements from all solutions
</output>
Best Practices
Meta-Judge + Judge Verification
- Two meta-judges - Separate specs for pruning (proposals) and evaluation (full solutions)
- Meta-judges run in parallel with implementation - Don't block the pipeline; pruning meta-judge runs with Phase 1, evaluation meta-judge runs with Phase 3
- Include CLAUDE_PLUGIN_ROOT - Both meta-judges and judges need the resolved plugin root path
- Meta-judge YAML - Pass only the YAML to judges, do not modify it
Common Pitfalls
- Insufficient exploration - Agents propose similar approaches
- Ignoring judge feedback - Expansion ignores concerns from pruning
- Vague proposals - Can't properly evaluate without implementation details
- Over-exploration - Too many proposals, evaluation becomes expensive
- Forcing synthesis when clear winner exists - Wastes cost and risks degrading quality
- Synthesizing fundamentally flawed solutions - Better to redesign than polish garbage
Recommendations
- Encourage diverse exploration - Prompt for different regions of solution space
- Feed feedback forward - Expansion agents address pruning concerns
- Right level of detail - Proposals have enough detail to evaluate
- Prune aggressively - Only expand most promising 3 approaches
- Trust adaptive strategy selection - Polish clear winners, synthesize split decisions, redesign failures
Example: API Design
Phase 1 outputs (assuming date 2025-01-15):
.specs/research/users-api-2025-01-15.proposals.a.md- 6 approaches from Agent A.specs/research/users-api-2025-01-15.proposals.b.md- 6 approaches from Agent B.specs/research/users-api-2025-01-15.proposals.c.md- 6 approaches from Agent C
Phase 1.5 output (runs in parallel with Phase 1):
- Pruning Meta-judge (Opus,
sadd:meta-judge) generates pruning evaluation specification YAML
Phase 2 outputs (3 judges with pruning meta-judge spec):
.specs/research/users-api-2025-01-15.pruning.1.md- Top 3: Resource-based REST, Pure REST, Monolithic.specs/research/users-api-2025-01-15.pruning.2.md- Top 3: Pure REST, Hybrid (services), Resource-based REST.specs/research/users-api-2025-01-15.pruning.3.md- Top 3: Resource-based REST, REST+GraphQL hybrid, Pure REST.specs/research/users-api-2025-01-15.selection.md- Selected: Resource-based REST (8 pts), Pure REST (7 pts), Monolithic (4 pts)
Phase 3 outputs:
specs/api/users.a.md- Full resource-based design with nested routesspecs/api/users.b.md- Flat REST design with simple endpointsspecs/api/users.c.md- Monolithic API with service-oriented internals
Phase 3.5 output (runs in parallel with Phase 3):
- Evaluation Meta-judge (Opus,
sadd:meta-judge) generates evaluation specification YAML
Phase 4 outputs (3 judges with evaluation meta-judge spec):
-
.specs/reports/users-api-2025-01-15.1.md:"Prefers A for RESTfulness, criticizes C complexity"
-
.specs/reports/users-api-2025-01-15.2.md:"Prefers B for simplicity, criticizes A deep nesting"
-
.specs/reports/users-api-2025-01-15.3.md:"Prefers A for discoverability, criticizes B lack of structure"
Phase 4.5 decision (orchestrator parses headers):
- Split votes: A, B, A (no unanimous winner)
- Average scores: A=4.1, B=3.8, C=3.4 (all >=3.0)
- Strategy: FULL_SYNTHESIS
- Reason: Split decision with merit, synthesis needed
Phase 5 output (synthesis):
specs/api/users.md- Resource-based structure (from A), max 2-level nesting (from B), internal services (from C)
</output>


