
Eval Audit
ai-evals-course/evals-skills/skills/eval-audit作者 ai-evals-course80d5f7b0127c7572ed9e9339937adbfd7240ffeb無授權條款1.4K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫2 週前更新
Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT use when the goal is to build a new evaluator from scratch (use error-discovery, write-judge-prompt, or validate-evaluator instead).
- 80d5f7b0127c7572ed9e9339937adbfd7240ffeb目前提交 80d5f7b發布於 2026年10月9日
來源與署名
來源:ai-evals-course/evals-skills位於skills/eval-audit提交80d5f7b
授權條款: 無授權條款
內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。
更多來自 ai-evals-course/evals-skills 的技能

Validate Evaluator
ai-evals-course
透過資料切分、TPR/TNR 指標與偏差校正,以人工標註校準 LLM 評判器。

Generate Synthetic Data
ai-evals-course
使用基於維度的元組生成方式,為 LLM 管線評估建立多樣化的合成測試輸入。

Evaluate Rag
ai-evals-course
指導 RAG 檢索與生成品質評估,涵蓋指標、資料集建置與分塊最佳化。

Evals Start
ai-evals-course
將評估相關請求路由到同一個外掛中合適的專用技能。
更多AI & Agents技能

Discernment Nudge
anthropics
在實質回答後附加 2-3 個具體追問,協助使用者查核事實、推理與缺少的背景。

Ppt Template Creator
anthropics
將使用者的 PowerPoint 範本轉換為可重用的技能,用來產生品牌簡報。

Skill Development
anthropics
指導建立 Claude Code 外掛技能,涵蓋結構、描述、漸進式揭露與驗證。

Plugin Structure
anthropics
說明 Claude Code 外掛的結構、資訊清單與元件配置方式。

Command Development
anthropics
指導建立 Claude Code 斜線命令,涵蓋結構、YAML frontmatter、參數與外掛功能。

Claude Md Improver
anthropics
稽核儲存庫中的 CLAUDE.md 檔案、評估其品質,並在取得核准後套用針對性的改進。