Guides engineers working with AI agents through eval-first execution, task decomposition, model routing and cost tracking.
- What it does
- This skill provides an operating framework for engineering workflows where AI agents perform most implementation work and humans control quality and risk. It defines completion criteria before execution, breaks work into independently verifiable 15-minute units, routes tasks to model tiers by complexity, and measures results with capability and regression evals. It also covers session strategy, review priorities for AI-generated code, and per-task cost discipline.
- When to use it
- Use it when planning or running engineering work that is largely delegated to AI agents and needs structure for verification and risk control. It fits teams that want eval-first loops, clear task decomposition, and deliberate model-tier selection. It is also useful for reviewing AI-generated code and tracking model cost per task.
- Requirements
- Instructions only; no scripts or bundled assets. It assumes access to AI agent tooling and multiple model tiers (Haiku, Sonnet, Opus) for routing, plus the ability to run capability and regression evaluations.