GKE Reliability
Routing Note: To generate Kubernetes YAML manifests (
Deployment,StatefulSet,Service,ConfigMap,HTTPRoute,PodDisruptionBudget), opengke-manifest-generation/SKILL.md.
This reference covers high availability and reliability configuration for GKE clusters and workloads.
MCP Tools:
get_cluster,get_k8s_resource,describe_k8s_resource,apply_k8s_manifest,list_k8s_events
Golden Path Reliability Defaults
Workflows
1. Verify Cluster High Availability
- If
locationis a region (e.g.,us-central1), the control plane is regional - If
locationshas multiple entries, nodes span multiple zones
2. Pod Disruption Budgets (PDBs)
PDBs ensure minimum pod availability during voluntary disruptions (node upgrades, autoscaler scale-down).
Check existing PDBs:
Create PDB:
Every production Deployment with 2+ replicas should have a PDB.
3. Health Probes
Every production container should have liveness and readiness probes. Startup probes are recommended for slow-starting apps.
Check existing probes:
Recommended probe configuration:
- Readiness: Determines when a pod can accept traffic
- Liveness: Determines when to restart a container
- Startup: Disables liveness/readiness until the app is ready (prevents premature restarts)
4. Graceful Shutdown
Ensure applications handle SIGTERM and drain in-flight requests:
5. Topology Spread Constraints
Distribute pods across zones and nodes to survive failures:
- Zone spread (
DoNotSchedule): Hard requirement -- pods must be balanced across zones - Node spread (
ScheduleAnyway): Best-effort -- prefer distribution but don't block scheduling
6. Replicas
Best Practices & Production Guidelines
- Regional clusters for production: Always use regional clusters to survive zone failures.
- PDBs for everything: Every production workload with 2+ replicas needs a PodDisruptionBudget (PDB) to protect against voluntary disruptions.
- Probes with Explicit Timeouts: Every production container must have both
liveness and readiness probes defined. Always explicitly define
initialDelaySeconds,periodSeconds, andtimeoutSecondsfor all probes. Never rely on the Kubernetes default timeout of 1 second if your application requires more, but always set a strict limit to prevent hanging connections. - Zone spreading: Use topology spread constraints to distribute pods across failure domains (zones and nodes).
- Graceful shutdown: Handle
SIGTERMand set appropriateterminationGracePeriodSecondswith apreStopsleep hook to allow load balancer deregistration. - Maintenance windows: Schedule upgrades during low-traffic periods (see
the
gke-upgradesskill).

