
Deepeval
作者 confident-aic144abbce848Apache-2.018K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫昨天更新
DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app; add evals; generate datasets or goldens; use deepeval generate; use deepeval test run; send results to Confident AI; monitor production; run online evals; inspect traces; or iterate on prompts, tools, retrieval, or agent behavior from eval failures. AI agents are the primary use case. Covers Python SDK, pytest eval suites, CLI generation, traced evals, Confident AI reporting, and agent-driven improvement loops. DO NOT TRIGGER for unrelated generic pytest, non-AI test setup, or non-DeepEval observability work unless the user asks to compare or migrate to DeepEval; for instrumenting an app with DeepEval tracing, @observe, or framework integrations (use the `deepeval-tracing` skill); or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill).
僅公開檔案列表。將技能安裝到工作區後即可檢視檔案內容。
| 路徑 | 大小 | 類型 |
|---|---|---|
| LICENSE | 158 B | text/plain |
| references/artifact-contracts.md | 2.2 KB | text/markdown |
| references/choose-use-case.md | 1.6 KB | text/markdown |
| references/confident-ai.md | 3.8 KB | text/markdown |
| references/datasets.md | 2.6 KB | text/markdown |
| references/intake.md | 4.3 KB | text/markdown |
| references/iteration-loop.md | 4.1 KB | text/markdown |
| references/metrics.md | 9.3 KB | text/markdown |
| references/pytest-e2e-evals.md | 6.2 KB | text/markdown |
| references/synthetic-data.md | 9.1 KB | text/markdown |
| references/traced-evals.md | 2.7 KB | text/markdown |
| SKILL.md | 8.4 KB | text/markdown |
| templates/metrics.py | 1 KB | text/plain |
| templates/test_multi_turn_e2e.py | 717 B | text/plain |
| templates/test_single_turn_no_tracing.py | 917 B | text/plain |
| templates/test_single_turn_tracing.py | 537 B | text/plain |
來源與署名
來源:confident-ai/deepeval位於skills/deepeval提交c144abb
授權條款: Apache-2.0
內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。
更多來自 confident-ai/deepeval 的技能
更多AI & Agents技能

Skill Development
anthropics
指導建立 Claude Code 外掛技能,涵蓋結構、描述、漸進式揭露與驗證。

Plugin Structure
anthropics
說明 Claude Code 外掛的結構、資訊清單與元件配置方式。

Command Development
anthropics
指導建立 Claude Code 斜線命令,涵蓋結構、YAML frontmatter、參數與外掛功能。

Claude Md Improver
anthropics
稽核儲存庫中的 CLAUDE.md 檔案、評估其品質,並在取得核准後套用針對性的改進。

Smb Onboard
anthropics
引導小型企業主完成首次設定:連接工具、執行一次展現價值的配方、記錄業務背景並設定每週檢查節奏。

Smb Router
anthropics
將小型企業主的需求轉接到合適的外掛技能或指令,並說明可用功能。