Investigating Ci Failures

作者 PostHog469d1773e9cb無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Investigates a specific CI failure to a verdict: whose fault, which commit, who wrote it, and whether it's fixed. Use for "who broke master", "why did this test fail in CI", "is this failure my PR's fault or everyone's", "is this test flaky or actually broken", "when did this failure start". Works from the engineering_analytics warehouse views (engineering_analytics_ci_failures, engineering_analytics_ci_job_history) plus the CI failure logs. Not for aggregate CI health, cost, or merge bottlenecks (use diagnosing-ci-and-merge-bottlenecks) and not for building saved insights (use turning-engineering-analytics-into-insights).

AI 產生的概覽

調查某次具體的 CI 失敗並給出結論:是誰的問題、哪個提交、誰寫的,以及是否已修復。

功能
指導對單一失敗測試或紅色 CI 執行的調查,使用兩個 engineering_analytics 倉儲檢視和 CI 失敗日誌,無需重新執行 CI。它對失敗進行指紋識別,判斷其形態(單分支、合併佇列閘門、主幹中斷或偶發),並得出指明肇因提交、作者、PR 及修復情況的結論。它也列出注意事項,例如日誌僅含失敗、指紋僅支援 pytest,以及各資料來源的新鮮度差異。
適用情境
適用於諸如誰弄壞了主幹、某個測試為何在 CI 中失敗、失敗是該 PR 的問題還是主幹引入的、某個測試是偶發還是確實損壞,以及失敗從何時開始等問題。不適用於 CI 整體健康狀況、成本或合併瓶頸分析,也不用於建立已儲存的洞察。
執行需求
需要存取 engineering_analytics 倉儲檢視 engineering_analytics_ci_failures 和 engineering_analytics_ci_job_history,以及 CI 失敗日誌。它引用 MCP 工具(engineering-analytics-broken-tests、engineering-analytics-flaky-tests、engineering-analytics-ci-failure-logs、engineering-analytics-run-failure-logs)和一個 SQL 查詢參考檔案。不附帶指令碼。

Investigating CI failures

The job: take one failing test or one red run and get to a verdict a developer can act on — yours / trunk-borne / flaky, and when trunk-borne: the culprit SHA, its author, the PR, and whether a fix already landed. Everything below is derivation over data that already exists; you never need to re-run CI to answer.

Two warehouse views are the substrate (both non-materialized — always current, query them freely):

  • engineering_analytics_ci_failures — one row per pytest FAILED <nodeid> line from CI logs, pre-fingerprinted (fingerprint = test id + digit/hex-normalized error). Group by fingerprint to get first/last seen, occurrence count, and branch spread.
  • engineering_analytics_ci_job_history — one row per job attempt with conclusion AND commit attribution: head_sha, commit_author_name, commit_message, commit_pr_number (the merged PR that produced the commit, the only PR attribution a master push run has). This is where greens live; the logs are failure-only, so every "when did it turn red / green again" question must come from here, never from the logs.

Copy-ready SQL for every step is in references/investigation-queries.md.

Start wide: what's broken right now

For "what CI failures should I care about right now" (before you have a specific test in hand), the engineering-analytics-broken-tests MCP tool does the shape classification below across all live failures at once: it groups the last 2 days of failures by fingerprint and labels each breaking_master / blocking_merge_queue / novel_burst / potentially_resolved / flaky / pr_only, most urgent first, plus breaking_master_jobs (default-branch jobs whose latest run is red). Use it as the triage entry point, then drop into the per-failure workflow below to reach a culprit. It is the automated counterpart to fingerprinting by hand; the manual queries stay the way to pin a specific failure to a boundary and author.

blocking_merge_queue is the one shape the manual table below does not cover, because it looks like a single-branch failure and is not. The merge queue runs the full suite on a gate branch (trunk-merge/pr-<n>/…) carrying master, the PR, and every PR queued ahead of it with overlapping impacted targets (in practice most of the queue), so a failure there is on a commit that already passed the PR's own CI: a conflict with what landed or queued in between, or a queue-mate's own break, not that PR's bug by default. Read it as "this stopped a merge", list what the gate branch ran (query 8), and diff it against trunk rather than reading the PR alone.

The four failure shapes

Fingerprint the failure first (query 1 in the references), then read its shape — the classification falls out of three columns:

ShapeReadingNext step
1 branch, any windowThat PR's own problemRead its failure lines; done
1 trunk-merge/pr-<n>/… gate branchQueue-mate, or landed sinceQuery 8, then diff the gate branch against trunk
Many branches, dense burst, hits masterTrunk break (master is/was red)Boundary query → culprit (below)
Many branches, sporadic over days/weeksFlakyCorroborate with engineering-analytics-flaky-tests

Why cross-branch means trunk: PR CI runs the PR merged with master, so one bad master commit fails every concurrently-running PR. A failure appearing on many unrelated branches in a tight window is the signature of a master-merge break, not of those PRs' code. Tell the asker explicitly when their PR is not at fault — that is usually the single most valuable sentence in the answer.

Trunk break → culprit

Run the boundary query (query 2): master-only job history for the failing job, ordered by created_at. The pattern reads directly:

text
... success success | failure failure ... failure | success ...                    ^ first red = the culprit row  ^ first green = the fix row

The culprit row carries everything: head_sha, commit_author_name, commit_message (which names what changed), commit_pr_number. The first-green row identifies the fix the same way. Confidence check before naming anyone: does the culprit commit plausibly touch the failing area (its message / PR diff vs the failing test's module)? A boundary landing on an unrelated commit means sharding or timing noise — widen the window and check the adjacent commit before asserting.

Then verify the failure window in ci_failures matches (first_seen just after the culprit merged, last_seen shortly after the fix as the PR queue drained). Mismatch = you're looking at two different problems sharing a test.

Flaky → corroborate, don't guess

Read the sibling attempts before you say "flaky". One run cannot separate a flake from a deterministic failure, and ci_job_history already carries every attempt's conclusion for the branch. If every run on a branch is red, there is no flake to find, and offering a re-run wastes a run: a retry changes none of the inputs.

Sporadic shape alone is suggestive, not proof. The engineering-analytics-flaky-tests MCP tool reads per-test CI spans (rerun-pass signal — a test that failed then passed on retry in the same job) and is the stronger signal where it has coverage. Counts only, never rates: passing runs below the emitter's duration threshold aren't recorded, so there is no honest denominator.

Caveats you must carry into every answer

  • The logs are failure-only. No green baseline exists in ci_failures; absence of a fingerprint is weak evidence (the job may simply not have run). Greens come from ci_job_history only.
  • Fingerprints are pytest-only (v1). Jest / playwright / cargo failures appear in the raw failure logs but are not in ci_failures. For those, fall back to the raw failure logs via the engineering-analytics-ci-failure-logs (PR-scoped) / engineering-analytics-run-failure-logs (run-scoped) MCP tools.
  • A job that failed before its tests ran is invisible to every test-level surface. No FAILED line means no fingerprint, no span, nothing for the flaky-tests tool to read — the failure is real but the test-level answer is silence, not "fine". Setup, docker, and runner failures are only visible as job conclusions (query 7), which is also the one surface here that yields an honest rate, since it records greens.
  • Check runs-on before blaming a cache. Jobs on depot-* runners resolve actions/cache against Depot Cache, which the GitHub Actions cache API cannot see, so an empty result there is not evidence of eviction. Read the job's own Cache not found for input keys: line for the keys actually requested. Both steps stay green either way: actions/cache/restore succeeds on a miss, and a failed actions/cache/save is only a warning, so a green producer job may have stored nothing.
  • Freshness differs per source. Logs stream in near-real-time; the warehouse jobs/runs tables arrive via webhook sync and can lag. During a live incident, start from ci_failures and check the warehouse's max(created_at) before trusting a boundary (query 5). A boundary computed against a stale warehouse names the wrong commit.
  • A run's conclusion can be stale until the workflow_run webhook settles it (SPEC §7) — treat a very recent "failure-free" tail with suspicion.
  • Retries: run_attempt > 1 rows are the same job re-run. A failure that clears on attempt 2 is flake signal; one that fails through attempt 5+ is deterministic.
  • Reverts: a revert shows up as a new first-green (or first-red) commit whose commit_pr_number is the reverting PR — attribution follows the revert, not the original.
  • Time-bound every logs query. The failure-log stream is large; unbounded scans hit the read cap. 14 days covers almost every investigation.
  • Pair the warehouse twin too. A ci_job_history query windowed on created_at alone forces a full jobs scan — the parsed timestamp is a computed column the parquet scan can't prune on. Add a coarse created_at_raw >= '<YYYY-MM-DD>' string floor (a day below the window) alongside the precise created_at bound so the scan skips; created_at stays the exact filter.

Choosing a surface

QuestionUse
"What's broken across CI right now?"engineering-analytics-broken-tests MCP tool (triaged, classified)
Same, but from a terminalhogli ci:insights — these endpoints, scoped to the checkout's repo
"Why did MY PR's CI fail?"engineering-analytics-ci-failure-logs MCP tool (PR-scoped, grouped)
"Who broke master / when did X start?"The two views, workflow above
"Is X flaky?"Shape from ci_failures + the flaky-tests tool
"Is this setup/infra failure common?"Job conclusions, query 7 (no test rows exist for it)
"What's failing on master right now?"breaking_master rows + breaking_master_jobs from the broken-tests tool
"Is CI slow / expensive / PRs stuck?"The diagnosing-ci-and-merge-bottlenecks skill
"Save this as a dashboard/insight"The turning-engineering-analytics-into-insights skill

Output expectations

Lead with the verdict and the exoneration/blame in plain words ("not your PR — master was broken between 08:01 and 09:58 UTC by #68727; fixed by #68855"), then the evidence: the boundary rows, the fingerprint window, occurrence/branch counts. Name the author factually (they authored the culprit commit), never accusatorially — the commit message and PR link let the reader judge the change, and half the time the "culprit" was a reasonable change with an unmocked test dependency.

來源與署名

來源:PostHog/ai-plugin位於skills/investigating-ci-failures提交469d177

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 PostHog/ai-plugin 的技能

Writing Simplified Technical English

PostHog

套用 ASD-STE100 簡化技術英語規則,讓代理撰寫的文字語意明確、方便執行。

Writing & Content2026年10月8日

Working With Task Comments

PostHog

透過 PostHog MCP exec 調度器讀取並解讀 PostHog 任務、成品和畫布上的留言。

Productivity & Workflow2026年10月8日

Working With Skills

PostHog

指導代理使用 PostHog 的 skill-* MCP 工具來探索、讀取、建立、更新與重構技能。

AI & Agents2026年10月8日

Working With Scouts

PostHog

說明如何把監看工作委派給 PostHog Signals 偵察代理、處理其回報,並長期調校整個代理團隊的操作手冊。

AI & Agents2026年10月8日

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

待分類2026年10月8日

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

待分類2026年10月8日