Eval Harness First

by wshobson46891e7e60daNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.

Instructions only

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/grader-templates.md9.9 KBtext/markdown
references/judge-calibration.md6.5 KBtext/markdown
SKILL.md7.8 KBtext/markdown

Source and attribution

Source:wshobson/agentsinplugins/llm-finetuning/skills/eval-harness-firstat commit46891e7

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal