Incident Response (experimental)
Experimental: this skill uses
context: forkto delegate to theincident-investigatoragent. In some Claude Code contexts the forked subagent does not inherit the plugin's MCP tools, in which case the agent will stop and report rather than fall back to bash/curl. If that happens, use the inline alternatives:
/rootly:brief <incident>— stakeholder summary/rootly:status— service health/rootly:oncall— current responders
You are helping the user investigate and respond to a production incident. This runs in a forked context to keep incident data separate from the main coding session.
Workflow
1. Identify the Incident
If $ARGUMENTS contains an incident reference:
- If it is already a UUID (36-char hex with hyphens), use it directly with
mcp__rootly__getIncident. - If it looks like a sequential reference (
4460,#4460,INC-4460), resolve to UUID via MCP:- Normalize it to the exact incident number format
INC-<number>(for example,4460becomesINC-4460). - Call
mcp__rootly__list_incidentswithpage_size=100,page_number=1, andsort=-created_at. - Scan the returned
incidentsarray for an exactincident_numbermatch. If found, use the correspondingincident_idas the UUID. - If page 1 does not contain the match, use page 1's newest
incident_numberto estimate the likely page for the target incident number, then callmcp__rootly__list_incidentsfor that page and at most one adjacent page. - On every page, match only on
incidents[*].incident_numberand use the pairedincident_idwhen you find the exact match. - If the exact incident number is still not found quickly, stop and ask the user for the incident UUID instead of scanning indefinitely.
- Normalize it to the exact incident number format
- Use the resolved UUID for all subsequent MCP calls.
- Never use
mcp__rootly__search_incidentsfor numeric incident resolution, because that tool searches title/summary text rather than incident numbers. - Never walk paginated lists indefinitely. If the sequential number isn't found in the bounded lookup above, ask the user for the UUID.
If no incident ID provided:
- Call
mcp__rootly__search_incidentsfiltered to active status (started) - If no active incidents, report "No active incidents found" and stop
- If multiple active incidents, list them sorted by severity (critical first, then high, then medium, then low) and ask the user to select one
- For long lists, show critical/high severity first with a note about additional lower-severity incidents
2. Gather Full Context
Once you have the incident ID:
- Call
mcp__rootly__getIncidentto get the full incident record - Call
mcp__rootly__get_alert_by_short_idor search alerts for associated alert details and timeline - Call
mcp__rootly__find_related_incidentsto find historically similar incidents - Call
mcp__rootly__suggest_solutionsto get resolution recommendations - Call
mcp__rootly__get_oncall_handoff_summaryfor current team status
3. Present Response Brief
4. Human-in-the-Loop for Write Operations
CRITICAL: NEVER execute write operations automatically. Always present them as recommendations and wait for explicit user confirmation.
Write operations include:
updateIncident(changing severity, status, or any incident field)- Adding or removing responders
- Posting status updates
- Escalating incidents
- Any other mutation of Rootly data
When the user approves an action, execute it and report the result.
5. Error Handling
- MCP tool errors: Report the specific error message and suggest manual steps (e.g., "Check the Rootly dashboard directly")
- Low confidence results: If
find_related_incidentsreturns scores below 0.3, explicitly flag: "These matches are low confidence -- consider manual investigation" - Missing data: If any tool call returns empty results, note it and continue with available data rather than failing entirely


