Exploring Ai Failures

作者 PostHog469d1773e9cb無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Find where an AI/LLM application is failing in production and surface the failure patterns, working from real traces. Use when someone wants to understand what's going wrong with an AI feature, find and categorize failure modes, triage errors, or investigate quality issues (wrong answers, ignored instructions, hallucinations, tool misuse) — "what's failing in my agent", "surface error patterns", "why are the responses bad", "find the common failure modes", "what should I fix next". Covers scoping to one use case, finding failing traces by whichever signal fits the context (code errors, metric outliers, trace-type slices, manual review, existing-eval spikes, clustering), and reading them into a ranked failure taxonomy.

AI 產生的概覽

讀取正式環境中的 AI/LLM 追蹤記錄,找出並排序 AI 功能的失敗模式。

功能
引導代理先聚焦到單一使用情境,再依程式錯誤、指標離群值、追蹤類型切片、分層抽樣、評估指標突增或叢集等訊號選擇要開啟的追蹤記錄,然後完整閱讀約 20 至 30 筆追蹤。最終產出一份簡短、依頻率排序的具名失敗模式清單,每個模式附上一到兩個可解析的追蹤深層連結,依據的是實際閱讀內容而非錯誤計數或指標彙總。文件也涵蓋低資料量情境,可透過壓力測試或小型合成資料集處理,並把結果交回使用者決定後續方向。
適用情境
適用於需要了解 AI 或 LLM 功能出了什麼狀況的情況,例如回答錯誤、忽略指令、幻覺或工具誤用,或詢問下一步該修什麼。也適合排查品質問題,並把最主要的失敗模式轉化為評估。
執行需求
需要存取 PostHog 的 LLM 分析工具(query-llm-traces-list、query-llm-trace、execute-sql、llma-evaluation-list、generate-app-url)以及正式環境的追蹤資料。僅為說明文件,未附帶指令碼。

Exploring AI failures

The highest-value thing you can do with production AI traffic is look at where it fails and name the patterns. The catch: most failures are silent. The model returns a clean response — HTTP 200, no exception — that is wrong, off-topic, ignores an instruction, or misuses a tool. Those never raise an error, and they're usually the failures worth caring about.

So this skill is about finding failures (loud and silent), reading them, and grouping them into a ranked set of failure modes you can act on: fix a prompt, file a bug, prioritize work, or turn the top mode into an automatic eval (creating-online-evaluations).

Everything below serves one irreducible activity: reading real traces. The queries only tell you which traces to open — they are never the answer. If you report a list of problems without having opened traces, you've described the loud minority (the things that throw errors) and missed the job.

This is bottom-up: the failure modes emerge from real traces, not from a list of generic metrics decided in advance. For reading a single trace in depth, lean on exploring-llm-traces; for emergent grouping at high volume, exploring-llm-clusters.

Tools

ToolPurpose
posthog:query-llm-traces-listList candidate traces — filter by error, sort by a metric, scope by type
posthog:query-llm-traceRead a trace in full to see what actually went wrong
posthog:execute-sqlFind metric outliers, discover the trace taxonomy, count failure modes
posthog:llma-evaluation-listFind existing evals whose failures might reveal a new mode
posthog:generate-app-urlBuild a region- and project-qualified deep link to a trace or list

Detailed queries for each strategy below are in references/finding-traces.md [blocked]. The full $ai_* event schema (and the events vs ai_events split for heavy content like $ai_input/$ai_output_choices) lives in exploring-llm-traces/references/events-and-properties.md.

Work with the user

Collaborate on scope and priorities — not on whether to do the work. Narrow with the user up front: which feature or use case? have they already seen something bad? is there a signal to follow (a thumbs-down, a ticket, a metric that looks off)? Once it's scoped, go read traces and come back with coded failure modes — don't stop to ask permission before the reading; that reading is the core activity, not an optional follow-up to offer. When the user doesn't know what to look for, drive the loop below and explain the reasoning as you go; keep the teaching opt-in.

Step 1 — Scope to one use case

Apps have a taxonomy of trace types, and each fails differently — a support chat hallucinates policy, a summarizer drops key points, an agent loops or misuses a tool. Evaluating or analyzing them together averages the signal away. Pick one, then find its filter (a $ai_trace_id prefix, a feature property, a model). If the user isn't sure how their traffic splits, discover the taxonomy first (query in references/finding-traces.md [blocked]).

Step 2 — Pick which traces to read

These are ways to select which traces to open — not answers in themselves. The queryable ones (error counts, metric aggregates) tell you where to look; they are never the output. Choose by the context and signals you have, and combine them:

  • Code errors ($ai_is_error) — the cheapest sweep and the least representative signal: it only catches exceptions and API failures, not the silent quality failures that matter most. Use it to grab a few traces to read, not as a tally of "the problems." Slightly more useful for structured-output or tool-calling pipelines, where some failures do surface as parse/schema errors.
  • Metric outliers — sort by output/input tokens, message length, cost, or latency and open the extremes. Runaway length, truncation, context bloat, and loops cluster at the tails.
  • One trace-type slice — narrow to a single kind of request so the traces you read share a taxonomy.
  • Stratified sample — when you have no specific signal (the common case), pull a mixed batch across slices and outcomes and read it. This is the default, not the fallback.
  • Existing-eval spikes — when evals already run, a jump in an eval's failures points you at traces to read (llma-evaluation-list + execute-sql over the $ai_evaluation events).
  • Clustering — at high volume, let groupings emerge to pick representative traces to read; see exploring-llm-clusters.

The trap. It's tempting to GROUP BY error messages, produce a ranked table, and stop. That table is the loud minority — failures that raise an exception. The failures that matter for most AI products complete with HTTP 200 and only appear when a human reads the trace. A ranking built from error or metric counts you never opened is not the deliverable — it's a pointer to what to read next. If a query for silent failures comes back empty or awkward, that's a signal to read traces, not to give up and report the loud ones.

Step 3 — Read a batch (this is the job)

Open and actually read the traces you selected — plan on roughly 20–30 for a use case. Read each one with query-llm-trace, whose one required argument is traceId; the value to pass is the trace's id from the query-llm-traces-list result. This step is not optional, and nothing substitutes for it. You cannot find silent failures with GROUP BY or by grepping outputs for "refusal" / "sorry" language, because you don't yet know the patterns to search for — reading is how you discover them. A clever SQL proxy that returns nothing is not evidence the failures aren't there; it means you have to read.

For each trace, note in plain language what went wrong — and jot down the trace's earliest-event timestamp alongside the note (it's right there in the trace you just read, and in query-llm-traces-list's createdAt). That timestamp and the trace ID is all you need to build a resolvable deep link in Step 4, so capturing it now saves a second round-trip later.

When a trace fails in a chain, record the first thing that broke — the root failure usually causes the downstream symptoms, and fixing it clears them. Group the notes into a few named failure modes ("ignores the date filter", "invents a policy", "drops the second question"); a later pass can help cluster your notes, but review the groupings yourself. Keep reading until new traces stop turning up new modes (tens of traces, not thousands — stop when it goes quiet).

Step 4 — Rank, link, and hand back to the user

Rank the modes you found by reading, roughly by how often they showed up in your sample — a handful usually dominate. Present a short, ranked list of named failure modes. For each mode, include one or two example trace deep links on your own — don't wait to be asked, and don't make the user request them.

You read these traces, but you can misread one — a trace that looks like a hallucination may be correct in context, and some of what you flag will be you misunderstanding the trace, not a real failure. So don't present the list as settled fact. Give the user a couple of linked examples per mode, ask them to open the links, then ask which mode they want to focus on next.

(A list assembled from error messages or metric counts you never read is the loud subset, not this — go back to Step 3.)

When there's little to look at

If the use case is new or low-volume and you can't find enough failures: widen the time window or loosen the slice first; then stress-test with inputs that deliberately probe the constraints you care about (edge cases, long or ambiguous inputs, adversarial phrasing); or generate a small synthetic set across the dimensions that matter (request type × user scenario), run it through the system, and read those traces. Treat synthetic results as a bootstrap, not ground truth — they're unreliable for high-stakes or niche domains.

Constructing UI links

query-llm-trace returns a _posthogUrl for the trace it read, so hand that one back instead of composing your own. Build a link yourself only for a trace you have an id for but haven't opened, and then only with posthog:generate-app-url — never hand-write the host or the /project/<id>/ prefix. The url must be a canonical catalog template; pass concrete ids via params, never inline them into the path.

  • Traces list: generate-app-url {url: "/ai-observability/traces"} (then filter to your use case)
  • Single trace: generate-app-url {url: "/ai-observability/traces/{id}", params: {id: "<trace_id>"}}

No trace link carries a timestamp, so append ?timestamp=<url_encoded_timestamp> to whichever URL you hand back — the trace page reads it to resolve older traces, and neither tool can express it.

Generated links resolve to the correct region host and project prefix (e.g. https://us.posthog.com/project/<id>/ai-observability/traces/<trace_id>), so a user not already on the target project still lands in the right place.

Tips

  • Reading is the job, not the last step. Aggregates, error counts, and scores are clues for which traces to open — never a substitute. Read a first batch before reporting anything, and don't ask permission to do it.
  • Don't over-index on errors. $ai_is_error is the loudest but least interesting signal; the failures worth your time usually complete without one.
  • The finding strategies are a menu for picking traces to read, not a pipeline and not the answer. Pick by context, combine freely, and don't force an order.
  • One use case at a time. Different trace types have different failure taxonomies — mixing them blurs the result.
  • Frequency over completeness. The goal is the modes that happen most, not every conceivable failure.
  • The output is a ranked list of named failure modes from traces you read — that artifact is what makes the next step (fix, prioritize, or eval) obvious.
  • Hand back linked examples, then let the user steer. Don't stop at a categorical table. Give one or two resolvable trace links per mode unprompted, ask the user to eyeball a couple.

Related skills

  • creating-online-evaluations — turn a ranked failure mode into a continuously-running eval
  • exploring-llm-clusters — compare behavior across clusters instead of reading traces one by one
  • exploring-llm-traces — the trace-reading mechanics this skill leans on

來源與署名

來源:PostHog/ai-plugin位於skills/exploring-ai-failures提交469d177

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 PostHog/ai-plugin 的技能

Writing Simplified Technical English

PostHog

套用 ASD-STE100 簡化技術英語規則,讓代理撰寫的文字語意明確、方便執行。

Writing & Content2026年10月8日

Working With Task Comments

PostHog

透過 PostHog MCP exec 調度器讀取並解讀 PostHog 任務、成品和畫布上的留言。

Productivity & Workflow2026年10月8日

Working With Skills

PostHog

指導代理使用 PostHog 的 skill-* MCP 工具來探索、讀取、建立、更新與重構技能。

AI & Agents2026年10月8日

Working With Scouts

PostHog

說明如何把監看工作委派給 PostHog Signals 偵察代理、處理其回報,並長期調校整個代理團隊的操作手冊。

AI & Agents2026年10月8日

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

待分類2026年10月8日

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

待分類2026年10月8日