
Deepeval
by confident-aic144abbce848Apache-2.018K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today
DeepEval evaluation workflow for AI agents and LLM applications. TRIGGER when the user wants to evaluate or improve an AI agent, tool-using workflow, multi-turn chatbot, RAG pipeline, or LLM app; add evals; generate datasets or goldens; use deepeval generate; use deepeval test run; send results to Confident AI; monitor production; run online evals; inspect traces; or iterate on prompts, tools, retrieval, or agent behavior from eval failures. AI agents are the primary use case. Covers Python SDK, pytest eval suites, CLI generation, traced evals, Confident AI reporting, and agent-driven improvement loops. DO NOT TRIGGER for unrelated generic pytest, non-AI test setup, or non-DeepEval observability work unless the user asks to compare or migrate to DeepEval; for instrumenting an app with DeepEval tracing, @observe, or framework integrations (use the `deepeval-tracing` skill); or for raw OpenTelemetry / OTLP export without the deepeval package (use the `deepeval-otel` skill).
Add to a SourceWeft workspace
- Open the skill in your dashboard and add it to a workspace.
- Enable it for the chats that should use it.
This skill includes scripts. They run in your workspace sandbox when the skill is used — review the file list and source before adding it.
Add to SourceWeftYou will be asked to sign in first, then taken straight to this skill.
Ask your agent to install it
Paste this prompt into Claude Code, Codex, Cursor or another agent that can run commands — or into SourceWeft chat. The agent reads this skill's install guide, shows you its source, license and scripts, and installs it with the SourceWeft CLI once you agree.
Read https://sourceweft.com/skills/gh-confident-ai-deepeval-deepeval/install.md and install the skill it describes. Before installing, show me its source, license and whether it ships scripts, and wait for my OK. Ask me before changing anything else on my machine.Install it yourself from a terminal
For Claude Code, Codex, Cursor and other local agents. The SourceWeft CLI fetches the skill from its source repository at the commit scanned here, and verifies every file against the hashes recorded when the skill was scanned. If anything differs, nothing is written.
npx @sourceweft/cli skills install gh-confident-ai-deepeval-deepevalAdd --agent claude-code, codex, cursor or universal to choose which agent gets it (Claude Code by default).
Installed locally, this skill's scripts run on your machine, not in a sandbox. Read them first — the CLI asks before installing.
Upstream installer — not verified by SourceWeft
The open-source skills installer fetches the same pinned commit, but does not check the files against the hashes SourceWeft recorded.
npx skills add https://github.com/confident-ai/deepeval/tree/c144abbce848a6dfbd35bbdaaab49a62bb3fb7b6/skills/deepevalSource and attribution
Source:confident-ai/deepevalinskills/deepevalat commitc144abb
License: Apache-2.0
Content belongs to its original authors. SourceWeft indexes it from a public repository.
More in AI & Agents

Skill Development
anthropics
Guides creation of Claude Code plugin skills, covering structure, descriptions, progressive disclosure and validation.

Plugin Structure
anthropics
Guides the structure, manifest, and component layout of Claude Code plugins.

Command Development
anthropics
Guides creation of Claude Code slash commands, covering structure, YAML frontmatter, arguments and plugin features.

Claude Md Improver
anthropics
Audits CLAUDE.md files in a repository, scores their quality, and applies approved targeted improvements.

Google Cloud Solution Agentic Ai Data Science Workflow
Guides design of a multi-product agentic data science architecture on Google Cloud, from requirements to deployment and validation.

Google Cloud Solution Agentic Ai Borderless Data Lakehouse
Guides design of a borderless multicloud open data lakehouse on Google Cloud, from requirements discovery to deployment and validation.