Chaos Engineering
Design experiments that surface real weaknesses in production systems — without becoming outages. Most "chaos engineering" attempts skip steady-state measurement, define no abort criteria, and have no blast-radius bound. This skill enforces the discipline that makes chaos experiments safe and useful.
When to use
- Planning a chaos experiment (what to break, where, when, how to abort)
- Calculating blast radius before running the experiment
- Reviewing an existing experiment plan for safety
- Choosing a chaos tool (Chaos Toolkit / Chaos Mesh / Litmus / Gremlin / AWS FIS)
- Writing a chaos experiment postmortem
- Running a Game Day exercise
When NOT to use
- General incident response (use
incident-response) - Threat hunting / red-team (use
red-team,threat-detection) - Performance load testing (different goal — chaos is about failure modes, not capacity)
- Production debugging (chaos discovers weaknesses preemptively, not after-the-fact)
Core principle: chaos without abort criteria is an outage
The 4 Principles of Chaos Engineering (Netflix, 2016):
- Build a hypothesis around steady-state behavior. Not "what breaks?" but "X holds; will it still hold under fault Y?"
- Vary real-world events. Inject realistic failures: kill nodes, slow networks, lose cache, throttle dependencies.
- Run experiments in production. Staging never has the same failure modes. Start small.
- Automate experiments to run continuously. One-off chaos is a press release; continuous chaos is engineering.
Add a fifth: Define abort criteria up front. A chaos experiment with no abort criteria is an outage by another name.
Quick start
The 3 Python tools
All stdlib-only. Run with --help.
experiment_designer.py
Generates a structured experiment plan from inputs. Enforces the required sections (hypothesis, steady-state metric, blast radius, abort criteria, rollback).
Outputs a markdown plan with: hypothesis, steady-state, attack, magnitude, duration, blast radius, abort criteria, rollback procedure, monitoring dashboards, and learning question.
blast_radius_calculator.py
Computes the blast radius of a planned experiment. Given traffic share + user population + duration, calculates expected affected users, expected error budget burn, and a risk score.
Outputs:
- Expected affected users
- Error budget consumed (in minutes of error budget)
- Risk score: GREEN / YELLOW / RED
- Recommendation: PROCEED / REDUCE / ABORT
GREEN = <1% error budget; YELLOW = 1-10%; RED = >10%.
experiment_postmortem.py
Produces a structured postmortem from an experiment plan + results. Catches the common postmortem failure modes: no learning recorded, no follow-up actions, blame-laden language.
Outputs markdown with: summary, hypothesis (was it confirmed/refuted?), what we learned, what surprised us, follow-up actions with owners, and link to next experiment.
The 7 attack types (taxonomy)
Different attacks reveal different weaknesses. See references/attack_taxonomy.md for full detail.
Pick the attack that matches the hypothesis. "What happens if X is slow?" → latency. "What happens if X loses network?" → partition.
Tooling chooser
Decision rules:
- k8s-only stack + OSS → Chaos Mesh or Litmus (Litmus has bigger experiment library)
- Multi-cloud + OSS → Chaos Toolkit
- AWS-heavy + simple needs → AWS FIS
- Enterprise + audit/compliance → Gremlin
See references/tooling_landscape.md for trade-offs.
Workflows
Workflow 1: Design and run a single experiment
Workflow 2: Game Day exercise
Workflow 3: Continuous chaos (game days → daily)
Composition with other skills
This skill explicitly composes with two others in this library:
Anti-patterns
- No hypothesis — "let's break things" is sabotage, not engineering
- No steady-state metric — without a baseline, you can't tell if X broke
- No blast radius bound — full-prod experiment without limits = outage
- No abort criteria — see above; this is mandatory
- No on-call coverage — chaos without monitoring is unmonitored production
- Chaos in staging only — staging never has prod failure modes
- Chaos in dev — useless; dev has different failure modes from prod
- One-off chaos — single experiment is a press release; learning requires recurrence
- Blame-laden postmortem — record causes, not blame; teams stop running chaos otherwise
References
references/chaos_principles.md— the 4 principles, history, when to startreferences/experiment_design.md— hypothesis structure, steady-state metrics, abort criteriareferences/attack_taxonomy.md— 7 attack types with examples and toolingreferences/tooling_landscape.md— Chaos Toolkit / Mesh / Litmus / Gremlin / FIS / DIY
Slash command
/chaos-experiment — interactive experiment design wizard that runs all 3 tools.
Asset templates
assets/experiment_template.md— fill-in plan templateassets/postmortem_template.md— structured postmortem template
Verifiable success
A team using this skill should achieve:
- 100% of chaos experiments have a written hypothesis, abort criteria, and blast-radius calculation
- Blast radius for any single experiment never exceeds 10% of error budget
- Mean time between chaos experiments <14 days (continuous, not one-off)
- Each experiment produces ≥1 follow-up action that gets shipped
- No chaos experiment escalates to a customer-impacting incident in trailing 90 days


