Respond

作者 Rootly-AI-Labs65832aa6ff7a無授權條款收錄於 2026年10月8日更新於 2026年10月8日

[experimental] Investigate and respond to a production incident. Pulls context, finds similar past incidents, suggests solutions, and enables coordination. Forked-subagent flow may not have MCP access in all Claude Code contexts -- prefer /rootly:status + /rootly:brief for now.

僅含說明DevOps & Cloud
AI 產生的概覽

透過 Rootly MCP 工具調查生產事故,彙整脈絡、相似歷史事故與建議方案,產出回應簡報。

功能
這個實驗性技能會引導代理調查以 UUID 或序號標示的生產事故。它會解析事故、取得完整紀錄、告警詳情、歷史相似事故、建議解決方案與值班狀態,然後輸出結構化的事故回應簡報,包含時間軸、相關事故、依信心排序的方案與回應人員。嚴重性變更、回應人員調整、狀態更新與升級等寫入操作僅會以建議形式提出,必須取得使用者明確同意。
適用情境
適用於需要調查生產事故並產出彙整回應簡報的情境,例如值班回應或事故分流。也適合在採取行動前尋找相似的历史事故與候選解決方案。此技能說明其仍屬實驗性,且在某些環境中 fork 子代理流程可能無法存取 MCP,此時建議改用內嵌替代方式。
執行需求
需要 Rootly MCP 工具(mcprootly*)以及帶有 rootly:incident-investigator 代理的 fork 脈絡;不隨附指令碼。需要事故識別碼或可存取進行中的事故,寫入操作需使用者明確確認。

Incident Response (experimental)

Experimental: this skill uses context: fork to delegate to the incident-investigator agent. In some Claude Code contexts the forked subagent does not inherit the plugin's MCP tools, in which case the agent will stop and report rather than fall back to bash/curl. If that happens, use the inline alternatives:

  • /rootly:brief <incident> — stakeholder summary
  • /rootly:status — service health
  • /rootly:oncall — current responders

You are helping the user investigate and respond to a production incident. This runs in a forked context to keep incident data separate from the main coding session.

Workflow

1. Identify the Incident

If $ARGUMENTS contains an incident reference:

  1. If it is already a UUID (36-char hex with hyphens), use it directly with mcp__rootly__getIncident.
  2. If it looks like a sequential reference (4460, #4460, INC-4460), resolve to UUID via MCP:
    • Normalize it to the exact incident number format INC-<number> (for example, 4460 becomes INC-4460).
    • Call mcp__rootly__list_incidents with page_size=100, page_number=1, and sort=-created_at.
    • Scan the returned incidents array for an exact incident_number match. If found, use the corresponding incident_id as the UUID.
    • If page 1 does not contain the match, use page 1's newest incident_number to estimate the likely page for the target incident number, then call mcp__rootly__list_incidents for that page and at most one adjacent page.
    • On every page, match only on incidents[*].incident_number and use the paired incident_id when you find the exact match.
    • If the exact incident number is still not found quickly, stop and ask the user for the incident UUID instead of scanning indefinitely.
  3. Use the resolved UUID for all subsequent MCP calls.
  4. Never use mcp__rootly__search_incidents for numeric incident resolution, because that tool searches title/summary text rather than incident numbers.
  5. Never walk paginated lists indefinitely. If the sequential number isn't found in the bounded lookup above, ask the user for the UUID.

If no incident ID provided:

  1. Call mcp__rootly__search_incidents filtered to active status (started)
  2. If no active incidents, report "No active incidents found" and stop
  3. If multiple active incidents, list them sorted by severity (critical first, then high, then medium, then low) and ask the user to select one
  4. For long lists, show critical/high severity first with a note about additional lower-severity incidents

2. Gather Full Context

Once you have the incident ID:

  1. Call mcp__rootly__getIncident to get the full incident record
  2. Call mcp__rootly__get_alert_by_short_id or search alerts for associated alert details and timeline
  3. Call mcp__rootly__find_related_incidents to find historically similar incidents
  4. Call mcp__rootly__suggest_solutions to get resolution recommendations
  5. Call mcp__rootly__get_oncall_handoff_summary for current team status

3. Present Response Brief

## Incident Response Brief
### Summary**[Incident title]** (ID: [id])- **Status**: [status] | **Severity**: [severity]- **Started**: [time] ([duration] ago)- **Affected services**: [list]
### Timeline[Key events from alert and incident data, chronological]
### Related Historical Incidents[Top matches from find_related_incidents]- [Incident title] ([date]) - Confidence: [score] - Resolution: [what fixed it][If all scores < 0.3: "Low confidence matches -- manual investigation recommended"]
### Suggested Solutions[From suggest_solutions, ranked by confidence]1. [Solution] (confidence: [score], source: [incident/runbook])
### Current Responders & On-Call- **Assigned**: [responders]- **On-call**: [name] (since [time])- **Next handoff**: [time]
### Available ActionsThe following actions require your explicit approval:- Update severity- Add responder- Post status update- Escalate to next on-call

4. Human-in-the-Loop for Write Operations

CRITICAL: NEVER execute write operations automatically. Always present them as recommendations and wait for explicit user confirmation.

Write operations include:

  • updateIncident (changing severity, status, or any incident field)
  • Adding or removing responders
  • Posting status updates
  • Escalating incidents
  • Any other mutation of Rootly data

When the user approves an action, execute it and report the result.

5. Error Handling

  • MCP tool errors: Report the specific error message and suggest manual steps (e.g., "Check the Rootly dashboard directly")
  • Low confidence results: If find_related_incidents returns scores below 0.3, explicitly flag: "These matches are low confidence -- consider manual investigation"
  • Missing data: If any tool call returns empty results, note it and continue with available data rather than failing entirely

來源與署名

來源:Rootly-AI-Labs/rootly-claude-plugin位於skills/respond提交65832aa

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架