Gke Reliability

作者 google55b4e13eba6d無授權條款21K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Improves GKE workload reliability, using PDBs, health probes, and topology spread constraints. Use when configuring GKE workload reliability, setting up PDBs, or configuring GKE health probes (liveness, readiness, startup). Don't use for generating K8s YAML manifests (use gke-manifest-generation) or disaster recovery and cluster backups (use gke-backup-dr).

精選僅含說明DevOps & Cloud
AI 產生的概覽

透過 Pod 中斷預算、健康探針、優雅關閉與拓撲分布限制來提升 GKE 工作負載可靠性。

功能
這份參考技能說明如何為 GKE 叢集與工作負載設定高可用性。內容涵蓋驗證區域叢集可用性、建立 Pod 中斷預算、定義存活、就緒與啟動探針、處理優雅關閉、套用拓撲分布限制,以及選擇副本數量。它提供黃金路徑預設值、YAML 範例,以及用來檢查現有資源的 MCP 或 kubectl 指令。
適用情境
適用於設定 GKE 工作負載可靠性、設定 PDB,或定義 GKE 健康探針的情境。不適用於產生 Kubernetes YAML 資訊清單,也不適用於災難復原與叢集備份。
執行需求
需要存取 GKE 叢集,並使用所列的 MCP 工具,或以 gcloud 與 kubectl 作為備援。不包含指令碼,僅為說明性內容。

GKE Reliability

Routing Note: To generate Kubernetes YAML manifests (Deployment, StatefulSet, Service, ConfigMap, HTTPRoute, PodDisruptionBudget), open gke-manifest-generation/SKILL.md.

This reference covers high availability and reliability configuration for GKE clusters and workloads.

MCP Tools: get_cluster, get_k8s_resource, describe_k8s_resource, apply_k8s_manifest, list_k8s_events

Golden Path Reliability Defaults

SettingGolden Path ValueNotes
Cluster typeRegional (4 zones:Control plane replicated across
: : us-central1-a/b/c/f) : zones :
Upgrade strategySURGE (maxSurge: 1)Rolling upgrades with extra
: : : capacity :
Auto-repairtrueUnhealthy nodes replaced
: : : automatically :
Auto-upgradetrueNodes follow control plane
: : : version :
Release channelREGULARBalanced freshness and stability
Stateful HAEnabledLeader election for stateful
: : : workloads :

Workflows

1. Verify Cluster High Availability

# MCP (preferred)get_cluster(name="projects/<PROJECT>/locations/<REGION>/clusters/<CLUSTER>",  readMask="location,locations,nodePools.locations")
# gcloud fallbackgcloud container clusters describe <CLUSTER> --region <REGION> \  --format="json(location, locations)" \  --quiet
  • If location is a region (e.g., us-central1), the control plane is regional
  • If locations has multiple entries, nodes span multiple zones

2. Pod Disruption Budgets (PDBs)

PDBs ensure minimum pod availability during voluntary disruptions (node upgrades, autoscaler scale-down).

Check existing PDBs:

# MCP (preferred)get_k8s_resource(parent="...", resourceType="poddisruptionbudget")
# kubectl fallbackkubectl get pdb --all-namespaces

Create PDB:

yaml
apiVersion: policy/v1kind: PodDisruptionBudgetmetadata:  name: my-app-pdb  namespace: defaultspec:  minAvailable: 2       # Or use maxUnavailable: 1  selector:    matchLabels:      app: my-app

Every production Deployment with 2+ replicas should have a PDB.

3. Health Probes

Every production container should have liveness and readiness probes. Startup probes are recommended for slow-starting apps.

Check existing probes:

# MCP (preferred)describe_k8s_resource(parent="...", resourceType="deployment", name="<APP>", namespace="<NS>")
# kubectl fallbackkubectl get deployment <APP> -n <NS> -o yaml | grep -E "livenessProbe|readinessProbe|startupProbe"

Recommended probe configuration:

yaml
spec:  containers:  - name: app    livenessProbe:      httpGet:        path: /healthz        port: 8080      initialDelaySeconds: 15      periodSeconds: 10      timeoutSeconds: 2      failureThreshold: 3    readinessProbe:      httpGet:        path: /readyz        port: 8080      initialDelaySeconds: 5      periodSeconds: 5      timeoutSeconds: 2      failureThreshold: 3    startupProbe:             # For slow-starting apps      httpGet:        path: /healthz        port: 8080      initialDelaySeconds: 10      periodSeconds: 5      timeoutSeconds: 2      failureThreshold: 30    # 30 * 5s = 150s max startup time
  • Readiness: Determines when a pod can accept traffic
  • Liveness: Determines when to restart a container
  • Startup: Disables liveness/readiness until the app is ready (prevents premature restarts)

4. Graceful Shutdown

Ensure applications handle SIGTERM and drain in-flight requests:

yaml
spec:  terminationGracePeriodSeconds: 30    # Default; increase for long-running requests  containers:  - name: app    lifecycle:      preStop:        exec:          command: ["/bin/sh", "-c", "sleep 5"]  # Allow LB to deregister

5. Topology Spread Constraints

Distribute pods across zones and nodes to survive failures:

yaml
spec:  topologySpreadConstraints:  - maxSkew: 1    topologyKey: topology.kubernetes.io/zone    whenUnsatisfiable: DoNotSchedule    labelSelector:      matchLabels:        app: my-app  - maxSkew: 1    topologyKey: kubernetes.io/hostname    whenUnsatisfiable: ScheduleAnyway    labelSelector:      matchLabels:        app: my-app
  • Zone spread (DoNotSchedule): Hard requirement -- pods must be balanced across zones
  • Node spread (ScheduleAnyway): Best-effort -- prefer distribution but don't block scheduling

6. Replicas

Workload TypeMinimum ReplicasReason
Stateless web/API2Survive single pod/node
: : : failure :
Critical services3Survive zone failure with zone
: : : spread :
Stateful (databases)3 (with replication)Application-level quorum
Batch/jobs1Ephemeral by nature

Best Practices & Production Guidelines

  1. Regional clusters for production: Always use regional clusters to survive zone failures.
  2. PDBs for everything: Every production workload with 2+ replicas needs a PodDisruptionBudget (PDB) to protect against voluntary disruptions.
  3. Probes with Explicit Timeouts: Every production container must have both liveness and readiness probes defined. Always explicitly define initialDelaySeconds, periodSeconds, and timeoutSeconds for all probes. Never rely on the Kubernetes default timeout of 1 second if your application requires more, but always set a strict limit to prevent hanging connections.
  4. Zone spreading: Use topology spread constraints to distribute pods across failure domains (zones and nodes).
  5. Graceful shutdown: Handle SIGTERM and set appropriate terminationGracePeriodSeconds with a preStop sleep hook to allow load balancer deregistration.
  6. Maintenance windows: Schedule upgrades during low-traffic periods (see the gke-upgrades skill).

來源與署名

來源:google/skills位於skills/cloud/gke-reliability提交55b4e13

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 google/skills 的技能

Dpop Adoption

google

精選

指導為 Google OAuth 平台實作 OAuth 2.0 DPoP(RFC 9449)傳送方約束的更新權杖。

Security21K今天更新

Finding Google Skills

google

精選

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

待分類21K今天更新

Spanner Basics

google

精選

指導 Google Cloud Spanner 的執行個體與資料庫管理、結構定義設計、查詢與效能診斷。

Data & Analytics21K今天更新

Secops Triage

google

精選

引導 SOC 分析師對 Google SecOps 安全警示進行分診,從調查到結案或升級。

Security21K今天更新

Secops Investigate

google

精選

指導 SOC 分析師在 Google SecOps 中使用 UDM 查詢與時間軸進行深入的安全事件與實體調查。

Security21K今天更新

Secops Hunt

google

精選

指導在 Google SecOps 中使用 UDM 查詢、IoC 回溯、普遍性與異常分析進行主動威脅狩獵。

Security21K今天更新