Validate Evaluator

hamelsmu/evals-skills/skills/validate-evaluator

by hamelsmu22418da2bfb159f28a1b0dcf64e969e14ae56c99No licenseListed Oct 9, 2026Updated Oct 9, 2026

Calibrate an LLM judge against human labels using data splits, TPR/TNR, and bias correction. Use after writing a judge prompt (write-judge-prompt) when you need to verify alignment before trusting its outputs. Do NOT use for code-based evaluators (those are deterministic; test with standard unit tests).

Instructions onlyAI & Agents

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
SKILL.md8.7 KBtext/markdown

Source and attribution

Source:hamelsmu/evals-skillsinskills/validate-evaluatorat commit22418da

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal