Agent Evaluation

davila7/claude-code-templates/cli-tool/components/skills/ai-research/agent-evaluation

by davila78da17d671b6fNo license32K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

Instructions only

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
SKILL.md2 KBtext/markdown

Source and attribution

Source:davila7/claude-code-templatesincli-tool/components/skills/ai-research/agent-evaluationat commit8da17d6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal