Deepeval Tracing

作者 confident-aic144abbce848Apache-2.018K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Instrument an AI application with DeepEval's native tracing so its behavior is visible in Confident AI. TRIGGER when the user wants to add DeepEval tracing or @observe to an LLM app, agent, RAG pipeline, or chatbot; wire a framework, model-provider, or vector-database integration (LangGraph, LangChain, OpenAI Agents, LlamaIndex, Pydantic AI, CrewAI, and others); choose between a native integration and manual instrumentation; set span types, tags, or metadata; or send DeepEval-SDK traces to Confident AI's Observatory. DO NOT TRIGGER for building DeepEval pytest eval suites, datasets, goldens, metrics, or deepeval test run (use the `deepeval` skill), or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill). This skill is purely DeepEval-SDK instrumentation — producing well-formed traces, not running evals.

僅含說明AI & Agents
AI 產生的概覽

為 Python AI 應用程式接入 DeepEval 原生追蹤,讓 span 顯示在 Confident AI 的 Observatory 中。

功能
指導為 AI 應用程式加入 DeepEval SDK 追蹤:若有支援的框架、模型供應商或向量資料庫整合就優先採用,否則改用人工 @observe 埋點。內容涵蓋指定 span 型別(llm、retriever、tool、agent)、擷取輸入與輸出,以及加入 trace 層級的標籤與中繼資料。產出是可在 Confident AI 的 Observatory 中檢視的完整 trace,並不負責執行評估。
適用情境
適用於為以 Python 撰寫的 LLM 應用程式、代理、RAG 流程或聊天機器人接入 DeepEval 追蹤或 @observe。也適合在原生整合與人工埋點之間做選擇,或設定 span 型別、標籤與中繼資料。不適用於建立 DeepEval pytest 評估套件、資料集、指標,或原始 OpenTelemetry 匯出。
執行需求
需要 Python 並安裝 deepeval 套件(pip install deepeval);埋點使用 deepeval.tracing。將 trace 傳送到 Confident AI 需要 deepeval login 或匯出的 CONFIDENT_API_KEY,並需要網路存取。不附帶指令碼,僅為說明文件加兩份參考文件。

DeepEval Tracing

Use this skill to instrument an AI application — an LLM app, agent, RAG pipeline, or chatbot — with DeepEval's native tracing so its execution is visible span by span in Confident AI's Observatory. The work is: pick a supported integration when one exists, fall back to manual @observe otherwise, give each span a meaningful type, and add tags and metadata.

This skill stops at producing well-formed traces. Attaching evaluation metrics and running evals is the deepeval skill's job.

Scope: AI Applications Only

Instrument only the AI parts of the system — agent loops and planning, LLM calls, retrieval / vector search, and tool calls. The span types (llm, retriever, tool, agent) describe AI components. Do not trace non-AI software (web servers, CRUD backends, infrastructure). If the target has no LLM, agent, retrieval, or tool-calling component, this skill does not apply.

When to Use vs the deepeval and deepeval-otel Skills

  • This skill (deepeval-tracing) — instrument an app with the DeepEval SDK (@observe, framework integrations) so traces reach Confident AI.
  • deepeval skill — build pytest eval suites: datasets, metrics, traced evals, deepeval test run, iteration. It runs evals against an app this skill instrumented.
  • deepeval-otel skill — instrument with the vendor-neutral OpenTelemetry SDK instead of the DeepEval SDK (raw OTLP, including non-Python apps).

The three are complementary. If unsure between this skill and deepeval-otel: use this one when the app is Python and you want the DeepEval SDK; use deepeval-otel when you want raw OpenTelemetry or the app is not Python.

Prerequisites

  • An AI application in Python with pip install deepeval.
  • For traces to reach Confident AI: deepeval login, or an exported CONFIDENT_API_KEY (preferred for CI and non-interactive runs).

Workflow

  1. Confirm the target is an AI application (it has LLM calls, an agent loop, retrieval, or tool calls). If it has none of these, stop — this skill does not apply.
  2. Detect the framework, model provider, agent SDK, and vector database in use.
  3. Read references/integrations.md and the exact integration doc for what was detected. Prefer a native integration over manual instrumentation.
  4. If no native integration fits, instrument manually with @observe. Read references/tracing.md.
  5. Give each span a meaningful type (llm, retriever, tool, agent) and capture inputs/outputs.
  6. Add trace-level tags and metadata where they help diagnose failure patterns. Never trace secrets, credentials, or raw sensitive data.
  7. Confirm deepeval login or CONFIDENT_API_KEY, then verify traces appear in the Confident AI Observatory.

Core Principles

  1. Instrument AI components only — llm, retriever, tool, agent spans. Never trace non-AI software.
  2. Prefer a supported integration over manual @observe. Manual tracing is the fallback for unsupported frameworks and app-owned wrapper boundaries.
  3. Read the exact integration doc before writing tracing code.
  4. Give spans meaningful types; let names default to function names unless there is a strong reason to override.
  5. Never trace secrets, credentials, API keys, or raw sensitive user data.
  6. Producing traces is the scope. Attaching metrics and running evals belong to the deepeval skill; raw OpenTelemetry export belongs to deepeval-otel.

References

TopicFile
Manual instrumentation: @observe, span types, tags, metadatareferences/tracing.md
Integration selection rule and framework / model / vector-DB doc indexreferences/integrations.md

來源與署名

來源:confident-ai/deepeval位於skills/deepeval-tracing提交c144abb

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架