Judge

NeoLabHQ/context-engineering-kit/plugins/sadd/skills/judge

作者 NeoLabHQ23e2428e809d77717f8acc9659c374a3a1fcb93e无许可证收录于 2026年10月9日更新于 2026年10月9日

Launch a meta-judge then a judge sub-agent to evaluate results produced in the current conversation

仅含说明AI & Agents
AI 生成的概览

运行先元评审、后评审子代理的两阶段流程,为对话中先前产出的工作进行评分。

功能
该技能协调一条仅出报告的评估流程:先由元评审子代理生成针对性的评估规范,包含评分标准、检查清单和打分准则,再由评审子代理在隔离上下文中应用该规范。它会从对话中提取待评估的工作,派发两个子代理,校验返回的分数与证据,并呈现完整评估报告、结论和后续选项。产出为结构化评估报告,含各准则得分、证据引用和加权总分;不会修改被评估的工作。
适用场景
适用于需要对当前对话中已完成的工作进行独立、基于准则的质量评估时,例如审查代码改动、文档或分析。也适合在决定下一步之前,希望获得带证据引用的结构化评分的情形。
运行要求
仅为指令,不附带脚本。依赖代理的 Task 工具和指定的子代理类型(sadd:meta-judge 与 sadd:judge)、opus 模型设置,以及已解析的 CLAUDE_PLUGIN_ROOT 环境变量。

Judge Command

<task>

You are a coordinator launching a two-phase evaluation pipeline to assess work produced earlier in this conversation. First, a meta-judge generates tailored evaluation criteria. Then, a judge sub-agent applies those criteria with isolated context, structured scoring, and evidence-based feedback. The evaluation is report-only - findings are presented without automatic changes.

</task>

<context>

This command implements the meta-judge -> LLM-as-Judge pattern with context isolation:

  • Structured Evaluation: Meta-judge produces tailored rubrics, checklists, and scoring criteria before judging
  • Context Isolation: Judge operates with fresh context, preventing confirmation bias from accumulated session state
  • Evidence-Based: Every score requires specific citations from the work (file locations, line numbers)
  • Multi-Dimensional Rubric: Generated by meta-judge to match the specific artifact type and evaluation focus
  • Self-Verification: Dynamic verification questions with documented adjustments

</context>

Your Workflow

Phase 1: Context Extraction

Before launching the evaluation pipeline, identify what needs evaluation:

  1. Identify the work to evaluate:

    • Review conversation history for completed work
    • If arguments provided: Use them to focus on specific aspects
    • If unclear: Ask user "What work should I evaluate? (code changes, analysis, documentation, etc.)"
  2. Extract evaluation context:

    • Original task or request that prompted the work
    • The actual output/result produced
    • Files created or modified (with brief descriptions)
    • Any constraints, requirements, or acceptance criteria mentioned
    • Artifact type (code, documentation, configuration, etc.)
  3. Provide scope for user:

    Evaluation Scope:- Original request: [summary]- Work produced: [description]- Files involved: [list]- Artifact type: [code | documentation | configuration | etc.]- Evaluation focus: [from arguments or "general quality"]
    Launching meta-judge to generate evaluation criteria...

IMPORTANT: Pass only the extracted context to the sub-agents - not the entire conversation. This prevents context pollution and enables focused assessment.

Phase 2: Dispatch Meta-Judge

Launch a meta-judge agent to generate an evaluation specification tailored to the specific work being evaluated. The meta-judge will return an evaluation specification YAML containing rubrics, checklists, and scoring criteria.

Meta-Judge Prompt:

markdown
## Task
Generate an evaluation specification yaml for the following evaluation task. You will produce rubrics, checklists, and scoring criteria that a judge agent will use to evaluate the work.
CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`
## User Prompt{Original task or request that prompted the work}
## Context{Any relevant context about the work being evaluated}{Evaluation focus from arguments, or "General quality assessment"}
## Artifact Type{code | documentation | configuration | etc.}
## InstructionsReturn only the final evaluation specification YAML in your response.

Dispatch:

Use Task tool:  - description: "Meta-judge: Generate evaluation criteria for {brief work summary}"  - prompt: {meta-judge prompt}  - model: opus  - subagent_type: "sadd:meta-judge"

Wait for the meta-judge to complete before proceeding to Phase 3.

Phase 3: Dispatch Judge Agent

After the meta-judge completes, extract its evaluation specification YAML and dispatch the judge agent with both the work context and the specification.

CRITICAL: Provide to the judge the EXACT meta-judge evaluation specification YAML. Do not skip, add, modify, shorten, or summarize any text in it!

Judge Agent Prompt:

markdown
You are an Expert Judge evaluating the quality of work against an evaluation specification produced by the meta judge.
CLAUDE_PLUGIN_ROOT=`${CLAUDE_PLUGIN_ROOT}`
## Work Under Evaluation
[ORIGINAL TASK]{paste the original request/task}[/ORIGINAL TASK]
[WORK OUTPUT]{summary of what was created/modified}[/WORK OUTPUT]
[FILES INVOLVED]{list of files with brief descriptions}[/FILES INVOLVED]
## Evaluation Specification
```yaml{meta-judge's evaluation specification YAML}

Instructions

Follow your full judge process as defined in your agent instructions!

CRITICAL: You must reply with this exact structured evaluation report format in YAML at the START of your response!


CRITICAL: NEVER provide score threshold to judges in any format. Judge MUST not know what threshold for score is, in order to not be biased!!!
**Dispatch:**

Use Task tool:

  • description: "Judge: Evaluate {brief work summary}"
  • prompt: {judge prompt with exact meta-judge specification YAML}
  • model: opus
  • subagent_type: "sadd:judge"

### Phase 4: Process and Present Results
After receiving the judge's evaluation:
1. **Validate the evaluation**:   - Check that all criteria have scores in valid range (1-5)   - Verify each score has supporting justification with evidence   - Confirm weighted total calculation is correct   - Check for contradictions between justification and score   - Verify self-verification was completed with documented adjustments
2. **If validation fails**:   - Note the specific issue   - Request clarification or re-evaluation if needed
3. **Present results to user**:   - Display the full evaluation report   - Highlight the verdict and key findings   - Offer follow-up options:     - Address specific improvements     - Request clarification on any judgment     - Proceed with the work as-is
## Scoring Interpretation
| Score Range | Verdict | Interpretation | Recommendation ||-------------|---------|----------------|----------------|| 4.50 - 5.00 | EXCELLENT | Exceptional quality, exceeds expectations | Ready as-is || 4.00 - 4.49 | GOOD | Solid quality, meets professional standards | Minor improvements optional || 3.50 - 3.99 | ACCEPTABLE | Adequate but has room for improvement | Improvements recommended || 3.00 - 3.49 | NEEDS IMPROVEMENT | Below standard, requires work | Address issues before use || 1.00 - 2.99 | INSUFFICIENT | Does not meet basic requirements | Significant rework needed |
## Important Guidelines
1. **Meta-judge first**: Always generate evaluation specification before judging - never skip the meta-judge phase2. **Include CLAUDE_PLUGIN_ROOT**: Both meta-judge and judge need the resolved plugin root path3. **Meta-judge YAML**: Pass only the meta-judge YAML to the judge, do not modify it4. **Context Isolation**: Pass only relevant context to sub-agents - not the entire conversation5. **Justification First**: Always require evidence and reasoning BEFORE the score6. **Evidence-Based**: Every score must cite specific evidence (file paths, line numbers, quotes)7. **Bias Mitigation**: Explicitly warn against length bias, verbosity bias, and authority bias8. **Be Objective**: Base assessments on evidence and rubric definitions, not preferences9. **Be Specific**: Cite exact locations, not vague observations10. **Be Constructive**: Frame criticism as opportunities for improvement with impact context11. **Consider Context**: Account for stated constraints, complexity, and requirements12. **Report Confidence**: Lower confidence when evidence is ambiguous or criteria unclear13. **Single Judge**: This command uses one focused judge for context isolation
## Notes
- This is a **report-only** command - it evaluates but does not modify work- The meta-judge generates criteria tailored to the specific artifact type and evaluation focus- The judge operates with fresh context for unbiased assessment- Scores are calibrated to professional development standards- Low scores indicate improvement opportunities, not failures- Use the evaluation to inform next steps and iterations- Low confidence evaluations may warrant human review

来源与署名

来源:NeoLabHQ/context-engineering-kit位于plugins/sadd/skills/judge提交23e2428

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架