
Eval Audit
hamelsmu/evals-skills/skills/eval-auditby hamelsmu22418da2bfb159f28a1b0dcf64e969e14ae56c99No licenseListed Oct 9, 2026Updated Oct 9, 2026
Audit an LLM eval pipeline and surface problems: missing error analysis, unvalidated judges, vanity metrics, etc. Use when inheriting an eval system, when unsure whether evals are trustworthy, or as a starting point when no eval infrastructure exists. Do NOT use when the goal is to build a new evaluator from scratch (use error-analysis, write-judge-prompt, or validate-evaluator instead).
Add to a SourceWeft workspace
- Open the skill in your dashboard and add it to a workspace.
- Enable it for the chats that should use it.
This skill is instructions only: it ships no scripts to execute.
Add to SourceWeftYou will be asked to sign in first, then taken straight to this skill.
Ask your agent to install it
Paste this prompt into Claude Code, Codex, Cursor or another agent that can run commands — or into SourceWeft chat. The agent reads this skill's install guide, shows you its source, license and scripts, and installs it with the SourceWeft CLI once you agree.
Read https://sourceweft.com/skills/gh-hamelsmu-evals-skills-eval-audit-skills-eval-audit-319873c8863a068b/install.md and install the skill it describes. Before installing, show me its source, license and whether it ships scripts, and wait for my OK. Ask me before changing anything else on my machine.Install it yourself from a terminal
For Claude Code, Codex, Cursor and other local agents. The SourceWeft CLI fetches the skill from its source repository at the commit scanned here, and verifies every file against the hashes recorded when the skill was scanned. If anything differs, nothing is written.
npx @sourceweft/cli skills install @hamelsmu/eval-auditAdd --agent claude-code, codex, cursor or universal to choose which agent gets it (Claude Code by default).
Upstream installer — not verified by SourceWeft
The open-source skills installer fetches the same pinned commit, but does not check the files against the hashes SourceWeft recorded.
npx skills add https://github.com/hamelsmu/evals-skills/tree/22418da2bfb159f28a1b0dcf64e969e14ae56c99/skills/eval-auditSource and attribution
Source:hamelsmu/evals-skillsinskills/eval-auditat commit22418da
License: No license
Content belongs to its original authors. SourceWeft indexes it from a public repository.
More from hamelsmu/evals-skills

Validate Evaluator
hamelsmu
Calibrates an LLM judge against human labels using data splits, TPR/TNR metrics and bias correction.

Generate Synthetic Data
hamelsmu
Guides creation of diverse synthetic test queries for LLM pipeline evaluation using dimension-based tuple generation.

Evaluate Rag
hamelsmu
Guides evaluation of RAG retrieval and generation quality, including metrics, synthetic QA datasets, and chunking optimization.

Error Analysis
hamelsmu
Guides systematic error analysis of LLM pipeline traces to build a catalog of failure modes and prioritize fixes.
More in AI & Agents

Discernment Nudge
anthropics
Appends 2-3 specific follow-up questions to substantive answers so users can check facts, reasoning and missing context.

Ppt Template Creator
anthropics
Turns a user's PowerPoint template into a reusable skill that generates branded presentations.

Skill Development
anthropics
Guides creation of Claude Code plugin skills, covering structure, descriptions, progressive disclosure and validation.

Plugin Structure
anthropics
Guides the structure, manifest, and component layout of Claude Code plugins.

Command Development
anthropics
Guides creation of Claude Code slash commands, covering structure, YAML frontmatter, arguments and plugin features.

Claude Md Improver
anthropics
Audits CLAUDE.md files in a repository, scores their quality, and applies approved targeted improvements.