Eval Harness First

by wshobson46891e7e60daNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.

Instructions only
  1. 46891e7e60daCurrentcommit 46891e7Published Oct 8, 2026

Source and attribution

Source:wshobson/agentsinplugins/llm-finetuning/skills/eval-harness-firstat commit46891e7

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal