Validate Evaluator

ai-evals-course/evals-skills/skills/validate-evaluator

by ai-evals-course80d5f7b0127c7572ed9e9339937adbfd7240ffebNo license1.4K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated 2 weeks ago

Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you need to verify alignment before trusting its outputs. Do NOT use for code-based evaluators (those are deterministic; test with unit tests per `write-code-eval`).

Instructions onlyAI & Agents

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
agents/openai.yaml249 Bapplication/yaml
SKILL.md8.7 KBtext/markdown

Source and attribution

Source:ai-evals-course/evals-skillsinskills/validate-evaluatorat commit80d5f7b

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal