Eval Audit

hamelsmu/evals-skills/skills/eval-audit

by hamelsmu22418da2bfb159f28a1b0dcf64e969e14ae56c99No license1.6K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated 7 weeks ago

Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT use when the goal is to build a new evaluator from scratch (use error-analysis, write-judge-prompt, or validate-evaluator instead).

ArchivedInstructions onlyAI & Agents
  1. 22418da2bfb159f28a1b0dcf64e969e14ae56c99Currentcommit 22418daPublished Oct 9, 2026

Source and attribution

Source:hamelsmu/evals-skillsinskills/eval-auditat commit22418da

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal