Deepeval

作者 confident-aic144abbce848Apache-2.018K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库昨天更新

DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app; add evals; generate datasets or goldens; use deepeval generate; use deepeval test run; send results to Confident AI; monitor production; run online evals; inspect traces; or iterate on prompts, tools, retrieval, or agent behavior from eval failures. AI agents are the primary use case. Covers Python SDK, pytest eval suites, CLI generation, traced evals, Confident AI reporting, and agent-driven improvement loops. DO NOT TRIGGER for unrelated generic pytest, non-AI test setup, or non-DeepEval observability work unless the user asks to compare or migrate to DeepEval; for instrumenting an app with DeepEval tracing, @observe, or framework integrations (use the `deepeval-tracing` skill); or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill).

包含脚本AI & Agents

仅公开文件列表。将技能安装到工作区后即可查看文件内容。

路径大小类型
LICENSE158 Btext/plain
references/artifact-contracts.md2.2 KBtext/markdown
references/choose-use-case.md1.6 KBtext/markdown
references/confident-ai.md3.8 KBtext/markdown
references/datasets.md2.6 KBtext/markdown
references/intake.md4.3 KBtext/markdown
references/iteration-loop.md4.1 KBtext/markdown
references/metrics.md9.3 KBtext/markdown
references/pytest-e2e-evals.md6.2 KBtext/markdown
references/synthetic-data.md9.1 KBtext/markdown
references/traced-evals.md2.7 KBtext/markdown
SKILL.md8.4 KBtext/markdown
templates/metrics.py1 KBtext/plain
templates/test_multi_turn_e2e.py717 Btext/plain
templates/test_single_turn_no_tracing.py917 Btext/plain
templates/test_single_turn_tracing.py537 Btext/plain

来源与署名

来源:confident-ai/deepeval位于skills/deepeval提交c144abb

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架