
Deepeval
作者 confident-aic144abbce848Apache-2.018K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库昨天更新
DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app; add evals; generate datasets or goldens; use deepeval generate; use deepeval test run; send results to Confident AI; monitor production; run online evals; inspect traces; or iterate on prompts, tools, retrieval, or agent behavior from eval failures. AI agents are the primary use case. Covers Python SDK, pytest eval suites, CLI generation, traced evals, Confident AI reporting, and agent-driven improvement loops. DO NOT TRIGGER for unrelated generic pytest, non-AI test setup, or non-DeepEval observability work unless the user asks to compare or migrate to DeepEval; for instrumenting an app with DeepEval tracing, @observe, or framework integrations (use the `deepeval-tracing` skill); or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill).
- c144abbce848当前提交 c144abb发布于 2026年10月8日
来源与署名
来源:confident-ai/deepeval位于skills/deepeval提交c144abb
许可证: Apache-2.0
内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。
更多来自 confident-ai/deepeval 的技能
更多AI & Agents技能

Skill Development
anthropics
指导创建 Claude Code 插件技能,涵盖结构、描述、渐进式披露与验证。

Plugin Structure
anthropics
指导 Claude Code 插件的结构、清单与组件布局。

Command Development
anthropics
指导创建 Claude Code 斜杠命令,涵盖结构、YAML frontmatter、参数与插件功能。

Claude Md Improver
anthropics
审查仓库中的 CLAUDE.md 文件,评估其质量,并在获得批准后应用有针对性的改进。

Smb Onboard
anthropics
引导小微企业主完成首次设置:连接工具、运行一次体现价值的配方、记录业务背景并设定每周检查节奏。

Smb Router
anthropics
将小企业主的需求转接到合适的插件技能或命令,并说明可用功能。