Investigating Logs

作者 PostHog469d1773e9cb無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Investigate logs in a PostHog project: verify a service or deployment is healthy, explain an error spike, triage an incident, or understand what a log stream is saying. Use when the user asks to "check the logs", asks whether a service, deploy, release, or change is working or broke anything, asks why errors are up or what changed, or wants the root cause of failures visible in logs. Routes the logs MCP tools (services overview, pattern mining, before/after pattern diffing, bucketed counts, facets, raw rows) so investigations start from summaries instead of raw rows or hand-written SQL over the logs table.

AI 產生的概覽

調查 PostHog 專案日誌,用於驗證服務健康狀態、解釋錯誤暴增並進行事件排查。

功能
指導代理使用日誌 MCP 工具在 PostHog 專案中進行逐步收斂的日誌調查:服務總覽、模式探勘、前後模式差異比對、分桶計數、 facet 與原始資料列。它為變更後的健康驗證、暴增解釋、事件排查、陌生日誌流定位以及尋找已知目標規定了工作流程。產出為結論(健康、回歸或無法判定),並附上可疑項目、證據與注意事項。
適用情境
適用於使用者要求檢查日誌、詢問某個服務、部署、發行或變更是否正常或是否弄壞了什麼、詢問錯誤為何上升或發生了什麼變化,或想找出日誌中可見故障的根本原因時。不適用於建立或調整日誌警示、產品事件分析,或對日誌表手寫 SQL。
執行需求
需要存取 PostHog 專案及其日誌 MCP 工具(服務總覽、模式、模式差異、計數、計數範圍、走勢圖、facet 值、屬性、query-logs)。僅為說明文件,不附帶指令碼。

Investigating logs

Investigation is a narrowing problem: summarize before you read. One posthog:logs-patterns call compresses millions of lines into at most 200 templates, and one posthog:logs-patterns-diff call answers "what is different about now vs. before" directly. Raw rows (posthog:query-logs) are the last step of an investigation, never the first.

When to use this skill

  • "Check the logs" / "is service X healthy?" / "did my deploy (or model bump, config change, migration) break anything?"
  • "Why are errors up?" / "explain this spike" / incident triage — "what changed?"
  • "What is this service logging?" — orienting in an unfamiliar or noisy stream.
  • Finding the log evidence for a failure reported elsewhere (an alert, an error-tracking issue, a user complaint).

When not to use this skill

  • Creating or tuning log alerts — that's authoring-log-alerts.
  • Analytics over product events, persons, or insights — that's querying-posthog-data.
  • HogQL exposes a logs table via posthog:execute-sql, but do not investigate through it: hand-written SQL over logs routinely hits read-byte caps and re-derives what the tools below do in one cheap call. Reserve SQL for the rare case of joining log-derived facts with non-log data.

Tools

ToolJob
posthog:logs-services-createTop-25 services with log_count, error_count, error_rate, sparkline. Orientation.
posthog:logs-patternsMine one window's message templates, ordered by frequency. "What is this stream saying?"
posthog:logs-patterns-diffDiff templates between two windows: new / rate-shifted / gone. "What changed?"
posthog:logs-count / posthog:logs-count-rangesScalar and time-bucketed counts for a filter. Localize volume before pulling rows.
posthog:logs-sparkline-queryVolume over time broken down by severity or service (the one bucketed view with a breakdown).
posthog:logs-facet-values-createDistribution of severity/service (or a resource attribute) under a filter.
posthog:logs-attributes-list / posthog:logs-attribute-values-listDiscover attribute keys and values before building filters.
posthog:query-logsRaw rows. Endpoint of every drill-down, entry point of none.

Each tool's own description documents its parameters and response shape — read it before calling.

Pick the workflow by question shape

"Is it healthy?" — post-deploy / post-change verification

The user changed something (deploy, model bump, config, migration) and wants to know the logs still look right.

  1. Pin down the change time and the affected service(s). Ask if the user hasn't said; the diff is meaningless without a boundary.
  2. Orient with posthog:logs-services-create: is the service still logging at all, and what is its error_rate now? A service that went silent fails verification just as hard as one that started erroring.
  3. posthog:logs-patterns-diff with query.dateRange from the change time to now and baselineDateRange set to a comparable window just before the change, scoped to serviceNames. New error/fatal templates right after a change are the classic regression signature; large rate_ratio shifts on existing error templates are the second thing to check.
  4. Check volume continuity with posthog:logs-count-ranges spanning before and after the boundary: a rate discontinuity (crash loop, restart storm, silence) shows up here even when message content looks unchanged.
  5. Drill only the suspects: pivot each suspicious pattern to raw lines via its match_regex with posthog:query-logs.

A pass verdict needs all three: no new error templates, no large error rate_ratio shifts, and continuous volume. Say which windows you compared — "healthy" is only as strong as the baseline.

"Explain this spike"

  1. Localize it: posthog:logs-count-ranges over the user's window, then recurse into the dense bucket(s) — each bucket's date_from/date_to feeds the next call. Stop after 3–4 levels.
  2. Explain it: posthog:logs-patterns-diff with the spike as query.dateRange and the window just before as baselineDateRange. The top new and rate_shift entries are the explanation. Do not mine both windows separately and diff by hand — the diff is one call.

Incident triage — "what broke?"

posthog:logs-patterns-diff first: incident window vs. a known-good window just before (or omit the baseline for same-window-last-week). Suspects are new entries and the biggest rate_ratio shifts; pivot each to raw lines. If the failing service is unknown, find it first with posthog:logs-facet-values-create faceting service_name under severityLevels: ["error", "fatal"].

"What is this stream saying?" — unfamiliar service

posthog:logs-patterns over the last hour, scoped to the service. Scan templates by estimated_count and non-zero error share in severity_counts. Widen the window or add searchTerm only if the answer isn't there.

Known needle — a specific message, attribute, or person

When the target is already precise (an error string, a request id, a distinct_id), skip pattern mining: discover the right keys with posthog:logs-attributes-list / posthog:logs-attribute-values-list, size the result with posthog:logs-count, then pull rows with posthog:query-logs.

Rules that keep investigations honest and cheap

  • Scope serviceNames (or a resource-attribute filter) on every call once the target service is known. Unscoped calls scan the whole team's stream and starve the pattern sample budget.
  • posthog:query-logs requires an explicit query.dateRange — omitting it is a 400, not a default window.
  • Pattern counts are sampled estimates (sampled: true); templates rarer than ~1 in 10,000 rows can be invisible. Absence of a rare template is not evidence it stopped.
  • Before trusting a wall of new entries in a diff, check baseline.total_count — a tiny or empty baseline (logging only just started) makes everything look new.
  • severityLevels matches the six canonical lowercase buckets against severity_text exactly. Zero rows on a severity filter → check the stored values with posthog:logs-attribute-values-list { key: "severity_text" }.
  • Budget: one services call, at most one patterns-diff per window pair, 3–4 count-ranges levels, and query-logs only for confirmed suspects with limit ≤ 100.

Output

Lead with the verdict, then the evidence:

  • Verdict: healthy / regressed / inconclusive, with the windows compared.
  • Suspects (if any): template, classification (new / rate_shift), estimated counts or rate_ratio, services, and 1–2 sample raw lines.
  • What was checked and what wasn't: services covered, windows, and any sampling or baseline caveats that limit confidence.

The user should be able to act on the verdict without re-running the investigation.

Related skills

  • authoring-log-alerts — turn the check behind a verdict into a continuous, low-noise alert
  • exploring-apm-traces — follow a suspicious log line into the request traces around it

來源與署名

來源:PostHog/ai-plugin位於skills/investigating-logs提交469d177

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 PostHog/ai-plugin 的技能

Writing Simplified Technical English

PostHog

套用 ASD-STE100 簡化技術英語規則,讓代理撰寫的文字語意明確、方便執行。

Writing & Content2026年10月8日

Working With Task Comments

PostHog

透過 PostHog MCP exec 調度器讀取並解讀 PostHog 任務、成品和畫布上的留言。

Productivity & Workflow2026年10月8日

Working With Skills

PostHog

指導代理使用 PostHog 的 skill-* MCP 工具來探索、讀取、建立、更新與重構技能。

AI & Agents2026年10月8日

Working With Scouts

PostHog

說明如何把監看工作委派給 PostHog Signals 偵察代理、處理其回報,並長期調校整個代理團隊的操作手冊。

AI & Agents2026年10月8日

Validating And Publishing Canvases

PostHog

Validate and publish a canvas source project safely: the source-project shape, declared capabilities, reading the current version pointer, iterating on validation diagnostics, guarded publishing with expected_current_version_id, staging a draft build and promoting it, waiting out the queued build, and recovering from a 409 version_conflict or a 429 capacity limit without overwriting concurrent work. Use whenever a canvas edit is ready to save, a draft build is wanted, a canvas publish or build returns diagnostics or a conflict, or a task needs to understand canvas version history.

待分類2026年10月8日

Understanding Billing Usage

PostHog

Explains PostHog billing usage and spend from the customer's visible Billing MCP tools. Use when the user asks why usage or spend is high, which product or project is driving usage, what a usage type means, how to reduce usage, what changed over time, why they got a usage change alert, or whether a spike/drop alert was real or noisy. Also use before product-specific analytics skills when the user names a billable PostHog product metric such as events, recordings, feature flag requests, exceptions, survey responses, synced rows, logs, AI events, AI credits, or Inbox credits. Starts from Billing usage/spend tools, then routes to customer-visible product MCP surfaces for deeper investigation.

待分類2026年10月8日