Agent Evaluation

davila7/claude-code-templates/cli-tool/components/skills/ai-research/agent-evaluation

by davila78da17d671b6fNo license32K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

Instructions only

Add to a SourceWeft workspace

  1. Open the skill in your dashboard and add it to a workspace.
  2. Enable it for the chats that should use it.

This skill is instructions only: it ships no scripts to execute.

Add to SourceWeft

You will be asked to sign in first, then taken straight to this skill.

Ask your agent to install it

Paste this prompt into Claude Code, Codex, Cursor or another agent that can run commands — or into SourceWeft chat. The agent reads this skill's install guide, shows you its source, license and scripts, and installs it with the SourceWeft CLI once you agree.

Read https://sourceweft.com/skills/gh-davila7-claude-code-templates-agent-evaluation/install.md and install the skill it describes. Before installing, show me its source, license and whether it ships scripts, and wait for my OK. Ask me before changing anything else on my machine.

Read the install guide the agent follows

Install it yourself from a terminal

For Claude Code, Codex, Cursor and other local agents. The SourceWeft CLI fetches the skill from its source repository at the commit scanned here, and verifies every file against the hashes recorded when the skill was scanned. If anything differs, nothing is written.

npx @sourceweft/cli skills install @davila7/agent-evaluation

Add --agent claude-code, codex, cursor or universal to choose which agent gets it (Claude Code by default).

Upstream installer — not verified by SourceWeft

The open-source skills installer fetches the same pinned commit, but does not check the files against the hashes SourceWeft recorded.

npx skills add https://github.com/davila7/claude-code-templates/tree/8da17d671b6f371d1b6a10b8309a4f8d101af8e5/cli-tool/components/skills/ai-research/agent-evaluation

Source and attribution

Source:davila7/claude-code-templatesincli-tool/components/skills/ai-research/agent-evaluationat commit8da17d6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal