Agent Evaluation

davila7/claude-code-templates/cli-tool/components/skills/ai-research/agent-evaluation

by davila78da17d671b6fNo license32K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Testing and benchmarking LLM agents including behavioral testing, capability assessment, reliability metrics, and production monitoring—where even top agents achieve less than 50% on real-world benchmarks Use when: agent testing, agent evaluation, benchmark agents, agent reliability, test agent.

Instructions only
  1. 8da17d671b6fCurrentcommit 8da17d6Published Oct 8, 2026

Source and attribution

Source:davila7/claude-code-templatesincli-tool/components/skills/ai-research/agent-evaluationat commit8da17d6

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal