Alert Investigation

作者 launchdarkly2fc544d3140fApache-2.026 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫昨天更新

Investigates a triggered observability alert and returns a structured diagnosis with likely cause, scope, and next steps.

AI 產生的概覽

調查已觸發的可觀測性警示,並回傳包含可能原因、影響範圍與後續步驟的結構化診斷。

功能
引導代理運用警示的結構化脈絡(例如警示 ID、名稱、閾值、越界值與時間範圍)調查特定的已觸發警示。它會依產品類型載入對應的日誌、追蹤、錯誤、工作階段或指標companion檔案,並在這些訊號上執行限定範圍的調查。最終產出結構化診斷,涵蓋觸發內容、可能原因、影響範圍與後續步驟,並引用追蹤 ID、錯誤群組 ID、工作階段 ID 與日誌時間戳等識別碼。
適用情境
適用於可觀測性警示已觸發,而值班人員或負責人需要了解觸發原因、受影響範圍以及下一步行動時。適合與日誌、追蹤、錯誤、工作階段或指標相關的警示,包括跨越產品邊界的複合警示。
執行需求
需要遠端託管的 LaunchDarkly 可觀測性 MCP 伺服器,以及用於查詢日誌、追蹤、錯誤群組、工作階段、聚合指標與屬性鍵的工具。不附帶指令碼,僅為指示文件。

Alert investigation

You are investigating a specific triggered alert. Alerts arrive with structured context — an alert ID, name, threshold, value that crossed it, and a time range. Your job is to explain why it fired, assess scope, and recommend action.

Prerequisites

This skill uses the following LaunchDarkly observability MCP tools:

  • query-logs — query log records
  • query-traces — query distributed traces
  • query-error-groups — query error groups
  • query-sessions — query sessions
  • query-aggregations — query aggregated/time-bucketed metrics
  • get-keys — discover available attribute keys before filtering

Workflow

  1. Parse the alert context. The first turn of the conversation carries alert variables: alertID, alertName, alertValue, group, groupValue, query, thresholdWindow, timeRange, plus a product-specific link. Use these, don't re-derive them.
  2. Load the per-product companion. Based on the alert's product type, load the matching companion: logs.md, traces.md, errors.md, sessions.md, or metrics.md. Each captures the per-product investigation shape.
  3. Run the investigation using the methodology from the investigate skill (cross-reference logs/traces/errors/sessions/metrics; cite identifiers; aggregate before paginating). Scoped to the alert's time range and filter.
  4. Produce a structured diagnosis. See output template below.

Output template

Alert investigations have a consistent structure so consumers (notification channels, dashboards) can parse them.

## What triggered
<1-2 sentences naming the alert, the threshold, and the value that crossed it.>
## Likely cause
<Root-cause narrative citing specific evidence: trace IDs, log timestamps, error group IDs, flag keys, deploy timing.>
## Scope
<Who or what is affected. Number of users, services, sessions, error groups. Time window of impact.>
## Next steps
<1-3 concrete actions the on-call or owner should take. Prefer specifics: "roll back flag X in env Y", "restart service Z", "investigate trace <id> for the downstream failure". Avoid "investigate further" — if you don't have a root cause, say what specifically should be investigated and how.>

When to load which companion

  • logs.md — log alert, log pattern alert
  • traces.md — latency alert, trace-error-rate alert, span-specific alert
  • errors.md — error-rate alert, new-error-group alert, crash-rate alert
  • sessions.md — session-health alert, user-facing-error-rate alert
  • metrics.md — custom metric threshold, aggregated metric alert, composite alert

If the alert crosses product boundaries (e.g. a metric alert driven by error data), load both companions.

Guidelines

  • Stay tight. Alert investigations feed notifications — keep the output structured and scannable. No preamble ("Here is my analysis..."), no repeated framing.
  • Cite identifiers. Every claim in the diagnosis should reference a specific trace ID, error group ID, session ID, or log timestamp.
  • If the alert appears to be noise, say so explicitly — "This alert fired because of <X>, but the underlying behavior is within normal variance because <Y>". Noise is a legitimate outcome; don't invent root causes.
  • Don't redo the investigation you just did. The diagnosis output should let the on-call act without re-querying.

來源與署名

來源:launchdarkly/ai-tooling位於skills/observability/alert-investigation提交2fc544d

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架