
Eval Audit
hamelsmu/evals-skills/skills/eval-audit作者 hamelsmu22418da2bfb159f28a1b0dcf64e969e14ae56c99无许可证1.6K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库7周前更新
Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT use when the goal is to build a new evaluator from scratch (use error-analysis, write-judge-prompt, or validate-evaluator instead).
仅公开文件列表。将技能安装到工作区后即可查看文件内容。
| 路径 | 大小 | 类型 |
|---|---|---|
| SKILL.md | 9.7 KB | text/markdown |
更多来自 hamelsmu/evals-skills 的技能

Validate Evaluator
hamelsmu
使用数据划分、TPR/TNR 指标和偏差校正,将 LLM 评审器与人工标注进行校准。

Generate Synthetic Data
hamelsmu
指导使用基于维度的元组生成方法,为 LLM 流水线评估创建多样化的合成测试查询。

Evaluate Rag
hamelsmu
指导 RAG 检索与生成质量评估,涵盖指标、合成问答数据集与分块优化。

Error Analysis
hamelsmu
指导对 LLM 流水线 trace 进行系统性错误分析,建立失败模式目录并确定修复优先级。
更多AI & Agents技能

Discernment Nudge
anthropics
在实质性回答后附加2-3个具体追问,帮助用户核查事实、推理与缺失背景。

Ppt Template Creator
anthropics
将用户的 PowerPoint 模板转化为可复用技能,用于生成品牌演示文稿。

Skill Development
anthropics
指导创建 Claude Code 插件技能,涵盖结构、描述、渐进式披露与验证。

Plugin Structure
anthropics
指导 Claude Code 插件的结构、清单与组件布局。

Command Development
anthropics
指导创建 Claude Code 斜杠命令,涵盖结构、YAML frontmatter、参数与插件功能。

Claude Md Improver
anthropics
审查仓库中的 CLAUDE.md 文件,评估其质量,并在获得批准后应用有针对性的改进。