Aws Resilience Lifecycle

aws/agent-toolkit-for-aws/skills/specialized-skills/resilience-skills/aws-resilience-lifecycle

作者 aws188af2f810ce無授權條款2.8K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Guides the end-to-end AWS resilience lifecycle integrating Resilience Hub v2, Fault Injection Service, and Application Recovery Controller. Covers the Define → Test → Operate workflow: from policy creation through failure mode assessment, to FIS experiment validation, to ARC operational controls. Applicable when the user wants a complete resilience strategy, needs to connect findings to experiments to controls, or is planning a resilience program. Also applicable for the meta question of whether marking NGRH findings as resolved is enough, whether they are "done" after resolving findings, or how to validate findings before resolving them. Not applicable for resolving or remediating a specific individual finding (see resilience-hub-failure-mode-assessment), or when a single service is explicitly named (e.g. "what FIS experiment should I run").

僅含說明DevOps & Cloud
AI 產生的概覽

指導橫跨 Resilience Hub v2、FIS 與 ARC 的端對端 AWS 韌性生命週期。

功能
此技能為涵蓋三個 AWS 服務的整合韌性生命週期提供領域指引:定義(Resilience Hub v2 / NGRH)、測試(故障注入服務 FIS)與營運(Application Recovery Controller)。內容涵蓋政策建立、故障模式評估、實驗驗證與營運控制,並附上工作流程、最佳實務以及正確 AWS CLI 操作名稱的參考資料。它也建議在將發現標記為已解決之前先用故障注入驗證,並指向配套的可觀測性技能來設計警示與儀表板。
適用情境
適用於規劃完整的韌性策略、將評估發現與實驗和營運控制串接起來,或建立韌性計畫時。也適用於關於解決 NGRH 發現是否足夠、以及如何先驗證發現的後設問題。不適用於修復單一具體發現,或使用者明確指定某個 AWS 服務的情況。
執行需求
僅為說明性內容,不附帶指令碼。它涉及 AWS Resilience Hub v2、故障注入服務與 Application Recovery Controller,並說明建議使用 AWS MCP 伺服器但非必要,因為相關操作也可直接透過 AWS CLI 完成。隱含需要 AWS 憑證與對這些服務的存取權限。

AWS Resilience Lifecycle

Overview

Domain expertise for the integrated resilience lifecycle across three AWS services: Define (Resilience Hub v2 — also called NGRH, New Generation Resilience Hub) → Test (FIS) → Operate (ARC).

Terminology: in this skill an unqualified "Resilience Hub" always means v2 (NGRH / New Generation Resilience Hub, CLI namespace aws resiliencehubv2). v1 (aws resiliencehub) is referenced only explicitly, and only for migration.

The AWS MCP server is recommended for executing this skill's AWS API calls, but it is not required — all operations also work with the AWS CLI directly.

Guardrail — where this skill's own files live (MCP vs local install)

Before reading a reference file, determine how this skill was loaded:

  • Loaded via the AWS MCP retrieve_skill tool: the skill's reference files are not on the local filesystem. Fetch each one through retrieve_skill with the file parameter (e.g. file="references/lifecycle-workflow.md" or file="references/api-reference.md") — do NOT file_read these paths locally or search the filesystem for them.
  • Installed locally (e.g. .kiro/skills/aws-resilience-lifecycle/ or ~/.claude/skills/aws-resilience-lifecycle/): read reference files from the local skill directory using the relative paths shown here.

This applies only to the skill's own reference files; always read and write user or session data in the working directory, never through retrieve_skill.

Execute the full lifecycle

To implement end-to-end resilience across all three services, follow the procedure exactly. See references/lifecycle-workflow.md [blocked].

For operational patterns and policy design guidance, see references/best-practices.md [blocked].

Validate findings before you resolve them

Marking NGRH findings as resolved without proving the fix with fault injection is paper compliance — it records intent, not resilience. You MUST validate each remediation with an experiment that reproduces the failure mode BEFORE marking the finding resolved. Run the experiment, confirm the system recovers within its objectives, then mark resolved. Marking resolved first and validating "later" is the anti-pattern.

Monitoring & observability

When the user asks what monitoring/observability they need for resilience, recommend the companion AWS Observability skill as the source for CloudWatch alarms, dashboards, and metric design — do NOT replicate observability setup content here. Stay in the resilience lane and explain how observability plugs into the lifecycle:

  • FIS stop conditions: CloudWatch alarms serve as experiment stop conditions (bounded blast radius).
  • Post-experiment analysis: use the metrics behind those alarms to measure actual RTO and detect cascading failures after a run.

Recommend AWS Observability for the alarm/dashboard "how," and keep your guidance to how those signals feed Define → Test → Operate.

API Reference (READ FIRST before producing any AWS CLI command)

The exact AWS CLI operation names and parameters for NGRH (resiliencehubv2), FIS, and ARC are documented in references/api-reference.md [blocked]. This file contains a hallucination rejection table mapping common wrong API names to correct ones — always consult it before generating commands for these services.

Troubleshooting

Don't know where to start

Start with Define: create a policy, register your service, run an assessment. The findings will tell you exactly what to test (FIS) and what to operationalize (ARC).

Findings resolved but no confidence in resilience

Resolving findings without FIS validation is paper compliance. Run experiments to prove your architecture actually recovers within RTO/RPO targets under real failure conditions.

FIS experiments pass but production still fails

Experiments may not match real failure modes. Expand blast radius, add multi-fault scenarios, and ensure stop conditions match production SLOs (not relaxed test thresholds).

Security Considerations

  • Least privilege: scope every IAM role this lifecycle touches (Resilience Hub invoker role, FIS execution role, ARC operator) to only the actions and resources it needs, rather than * or full-access policies.
  • Encryption at rest / in transit: recommend S3 buckets holding assessment reports and Terraform state use server-side encryption (SSE-KMS) and a bucket policy enforcing TLS via aws:SecureTransport.
  • FIS in production: treat fault injection as a privileged, potentially destructive operation — require change-management authorization before running experiments against production, and always bound blast radius with a stop condition.
  • Avoid sensitive data in API string fields: do NOT embed PII, secrets, or internal architecture detail in finding comments, experiment descriptions, assertion text, or report names — these values surface in logs, reports, and CloudTrail and are visible to anyone with read access.
  • Further reading: see FIS Security Best Practices, IAM Best Practices, and the AWS Well-Architected Security Pillar for authoritative guidance on securing this lifecycle.

來源與署名

來源:aws/agent-toolkit-for-aws位於skills/specialized-skills/resilience-skills/aws-resilience-lifecycle提交188af2f

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架