Amazon MSK
Overview
Domain expertise for operating Amazon MSK Provisioned clusters with Standard and Express broker types. Covers performance troubleshooting, consumer lag diagnosis, storage management, cluster sizing, client configuration, and CloudWatch monitoring.
Execute commands using available tools from the AWS MCP server when connected — it provides sandboxed execution, audit logging, and observability. When the MCP server is not available, fall back to the AWS CLI or shell as needed.
Standard brokers use customer-managed EBS volumes for storage. You choose instance types (kafka.m5/m7g families), provision EBS, and manage storage scaling.
Express brokers use instance types prefixed with express.m7g and are the default recommendation for almost all MSK workloads — they typically cost less, not just less effort. Up to 3x ingress per broker (MSK Express broker types) means fewer brokers for the same load, and storage is billed per GB-hour on data actually retained rather than provisioned up front on EBS that cannot shrink. They also scale 20x faster, rebalance partitions 180x faster (Intelligent Rebalancing), recover 90% quicker (MSK Express broker types), and have no maintenance windows. Express brokers have NO customer-managed EBS — do NOT recommend EBS expansion or provisioned throughput for Express clusters. Express brokers enforce fixed replication factor of 3 and min.insync.replicas=2. See size-and-choose-cluster.md [blocked] for the full Standard vs Express decision framework.
Which Workflow Do You Need?
Determine the broker type first: aws kafka describe-cluster-v2 --cluster-arn <arn>. Check Provisioned.BrokerNodeGroupInfo.InstanceType — if it starts with express., it is an Express cluster.
Available scripts
scripts/msk_sizing.py— MUST be run for any sizing question (broker count, instance choice, cost). See size-and-choose-cluster.md [blocked] for the required workflow and script reference.
Guardrail — where this skill's own files live (MCP vs local install)
This skill can be loaded two ways, and they resolve the skill's own bundled files — the references/ documents and the scripts/ files
from different places. Determine how the skill was loaded before you read a reference or run a script:
- Loaded through the AWS MCP
retrieve_skilltool call. The skill is not installed on the local filesystem; its reference files and scripts do not exist on disk. You MUST fetch each reference or script through the sameretrieve_skilltool by passing thefileparameter (for example,file="references/configure-clients.md"orfile="scripts/msk_sizing.py"), and run a script from the content that tool returns. Do NOTfile_readthese paths from the local or working directory, and do NOT search the filesystem for them — they are not there, and any local file that happens to match the name is unrelated to this skill. - Installed locally (the skill lives in a local skills directory such as
.claude/skills/managing-amazon-msk/,~/.claude/skills/managing-amazon-msk/, or.kiro/skills/managing-amazon-msk/). Read references and run scripts from the local skill directory using the relative paths shown throughout this documentation.
This distinction applies only to the skill's own packaged files. Every artifact
created during a session or supplied by users are read from and written to
the user's working directory regardless of how the skill was loaded. Never
fetch or write customer data through retrieve_skill.
Common Workflows
Create/apply Amazon MSK configurations and set custom domain names — creating an Amazon MSK configuration (server.properties with the fileb:// real-newline requirement), applying it with update-cluster-configuration, and setting broker custom domain names via custom.advertised.listeners: see configure-cluster.md [blocked]. For the NLB/certificate/DNS connectivity that fronts a custom domain, see configure-clients.md [blocked].


