Evals Start

ai-evals-course/evals-skills/skills/evals-start

by ai-evals-course80d5f7b0127c7572ed9e9339937adbfd7240ffebNo license1.4K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated 2 weeks ago

Entry point for evals. Use when the user asks for help with evals, does not know where to begin, or asks for something no other skill in this plugin matches. Do NOT use when a more specific skill in this plugin already matches; load that skill directly.

Instructions onlyAI & Agents
AI-generated overview

Routes eval-related requests to the appropriate specialized skill in the same plugin.

What it does
This skill acts as an entry point for evaluation work. It presents a table mapping user situations to specific skills, such as error discovery, eval auditing, code eval writing, judge prompt writing, evaluator validation, synthetic data generation, review interface building, and RAG evaluation. It tells the user which skill is being loaded and why, then hands off to that skill's workflow. It contains routing instructions only and produces no artifacts itself.
When to use it
Use when the user asks for help with evals, does not know where to begin, or requests something no other skill in the plugin matches. Do not use when a more specific skill in the plugin already matches; load that skill directly.
Requirements
No scripts or tools are required; it is an instructions-only routing document.

Evals Start

This plugin splits eval work into targeted skills. Your job here is small: find the row below that matches the user's situation, tell the user which skill you are loading and why, then load that skill and follow its workflow from start to finish instead of improvising your own version of it.

SituationSkill to load
Has traces, wants to find failure modes, no established taxonomy yeterror-discovery
Has an existing eval pipeline and wants to know if it can be trustedeval-audit
Has a known failure mode that code can check (e.g., format, schema, regex, execution)write-code-eval
Has a known failure mode and wants an LLM judge for itwrite-judge-prompt
Has an LLM judge or evaluator and wants to check its qualityvalidate-evaluator
Has no traces to review yetgenerate-synthetic-data, then error-discovery
Wants a custom annotation interface for some other labeling taskbuild-review-interface
Wants to evaluate a RAG pipelineevaluate-rag

Most requests that mention error analysis with traces in hand mean error-discovery. New users with an existing pipeline usually need eval-audit first. This file holds only routing. When in doubt about which row fits, ask the user instead of guessing. The workflow lives in the targeted skill.

Source and attribution

Source:ai-evals-course/evals-skillsinskills/evals-startat commit80d5f7b

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal