Gke Golden Path

作者 google55b4e13eba6d無授權條款21K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Provides GKE golden path configuration defaults, production readiness checklists, and cluster default patterns. Use when designing GKE clusters, verifying GKE production readiness, or checking configurations against GKE defaults. Don't use for setting up workload autoscaling specifically (use gke-workload-scaling instead).

精選僅含說明DevOps & Cloud
AI 產生的概覽

提供 GKE 黃金路徑叢集預設組態、生產就緒檢查清單與偏差指引,用於設計或稽核 GKE 叢集。

功能
提供生產 GKE 叢集建議的 Autopilot 組態預設值,包括一律套用的設定、可自訂設定及其取捨,以及 Day-0 與 Day-1 決策指引。它也提供生產就緒檢查、升級中斷指引,以及參考叢集政策 YAML 資產。稽核現有叢集時,會將其與黃金路徑比對,並回報偏差、嚴重程度與補救建議。
適用情境
適用於設計 GKE 叢集、驗證 GKE 生產就緒狀態,或對照 GKE 預設值檢查現有叢集組態。它並非專門用於設定工作負載自動擴縮,該主題由其他技能涵蓋。
執行需求
需要存取 GKE MCP 工具(get_cluster、create_cluster、update_cluster),或使用 gcloud 與 kubectl 進行叢集操作。必要輸入包括專案 ID、區域與叢集名稱。附帶參考 YAML 資產,但不含指令碼。

GKE Golden Path Configuration

The golden path is the recommended Autopilot configuration for production clusters. It defines sensible defaults — when the user requests different settings, apply them and note relevant trade-offs. For setting up autoscaling specifically, use gke-cluster-autoscaler for node autoscaling or gke-workload-scaling for workload autoscaling (HPA/VPA).

MCP Tools: get_cluster, create_cluster, update_cluster

Rules

  1. Default to the golden path. Use golden path values unless the user requests otherwise. When deviating, note trade-offs but respect the user's choice.
  2. Day-0 vs Day-1. Flag Day-0 decisions (networking, private nodes, subnets, IP allocation) prominently — they are hard/impossible to change after creation.
  3. Tool preference: MCP > gcloud > kubectl. MCP is preferred as it directly interfaces with GKE APIs with structured data, reducing shell syntax errors and parsing ambiguities. See the gke-basics skill's CLI reference for full coverage matrix and override options. If the user says "use gcloud" or "use kubectl", respect that for the session.
  4. Document decisions and rationale, especially for Day-0 choices and golden path deviations.

Required Inputs

If the user is unsure, use golden path defaults.

  • Project ID (required)
  • Region (required, e.g., us-central1)
  • Cluster name (required)
  • Environment type: dev/test or production (defaults to production)
  • Networking: bring-your-own VPC/subnet or auto-create (default: auto-create)
  • Scale expectations: expected node/pod count, workload types
  • Cost constraints: Spot VM tolerance, budget considerations

Always-Apply Defaults

Recommended best practices applied by default. If the user requests a different setting, apply it and briefly note the security or operational trade-off.

SettingGolden Path Value
autopilot.enabledtrue
privateClusterConfig.enablePrivateNodestrue
masterAuthorizedNetworksConfig.privateEndpointEnforcementEnabledtrue
secretManagerConfig.enabled + rotationInterval: 120strue
rbacBindingConfig.enableInsecureBinding*false (both)
workloadIdentityConfig.workloadPoolenabled
networkConfig.datapathProviderADVANCED_DATAPATH
networkConfig.dnsConfig.clusterDnsCLOUD_DNS
autoscaling.autoscalingProfileOPTIMIZE_UTILIZATION
verticalPodAutoscaling.enabledtrue
monitoringConfig componentsSYSTEM_COMPONENTS, STORAGE, POD, DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, JOBSET, CADVISOR, KUBELET, DCGM, APISERVER, SCHEDULER, CONTROLLER_MANAGER
loggingConfig componentsSYSTEM_COMPONENTS, WORKLOADS (enabled by default)
advancedDatapathObservabilityConfig.enableMetricstrue
nodeConfig.shieldedInstanceConfig.enableSecureBoottrue
nodeConfig.workloadMetadataConfig.modeGKE_METADATA
nodeConfig.gcfsConfig.enabled / gvnic.enabledtrue / true
addonsConfig.statefulHaConfig.enabledtrue
Storage CSI drivers (Filestore, GCS FUSE, Parallelstore)enabled
Pod Security Standardsrestricted on production namespaces

Customer-Configurable Settings

These have golden path defaults but customers may deviate with valid justification. Ask before changing.

SettingDefaultWhy Deviate
dnsEndpointConfig.allowExternalTraffictrueRestrict if cluster only accessed from within VPC
autoIpamConfig / createSubnetworktrue / trueCustomer has pre-existing VPC/subnets
maxPodsPerNode48 (this golden path's choice)Halves per-node IP consumption (/25 instead of /24). Not a GKE default (Standard defaults to 110, Autopilot to 32); raise for high pod-density at the cost of more CIDR space
subnetworkauto-createdCustomer brings existing subnets
Release channel + maintenance windowsREGULAR channel with a recurring maintenance windowAdd targeted maintenance exclusions (keep under ~6 months) only for critical freezes — see the gke-upgrades skill
nodeConfig.bootDisk.diskTypepd-balancedpd-ssd for I/O-intensive, pd-standard for cost

Note: Autopilot selects node machine types automatically (e.g., ek-standard-8 may appear in describe output); the machine type is not customer-configurable in Autopilot. Steer workload placement via ComputeClasses instead.

Guardrails

  • Do not request or output secrets (tokens, keys, service account JSON).
  • Resolve project/cluster context from the conversation, MCP tools, or gcloud config get-value project; ask the user only if it cannot be resolved.
  • For Day-0 decisions, always ask clarifying questions before proceeding.
  • For Day-1 features, propose golden path defaults with trade-offs and let the customer confirm.
  • Do not promise zero downtime — see Upgrade Disruption below for what to advise instead.
  • When auditing existing clusters, compare against golden path and report deviations with severity and remediation.

Upgrade Disruption

Never promise zero downtime for node upgrades, on any configuration. Node upgrades cordon and drain nodes, which evicts Pods. Draining honors PodDisruptionBudgets and terminationGracePeriodSeconds for up to one hour, after which GKE forcefully evicts the remaining Pods so the upgrade can proceed. A PDB narrows the window; it cannot veto the upgrade. Say so plainly rather than implying the disruption can be eliminated.

What to recommend, all four — not a subset:

  • PodDisruptionBudgets with minAvailable set so eviction cannot take the last healthy replica. A PDB that can never be satisfied stalls the drain for an hour and then loses anyway.

  • At least 2 replicas, spread across zones with topology spread constraints. A single-replica Deployment has downtime by definition.

  • Readiness probes that reflect real serving health, so traffic drains before the Pod dies.

  • Surge upgrade settings on the node pool. Surge is the default strategy; the default is maxSurge=1, maxUnavailable=0 — one extra node is created and made ready before an old one is drained.

    SettingControlsDefault
    maxSurgeAdditional nodes added per zone during the upgrade1
    maxUnavailableNodes simultaneously unavailable per zone0

    Nodes upgraded at once is the sum of the two, capped at 20 (Autopilot) and 100 (Standard). Multi-zone node pools upgrade one zone at a time. Raising maxUnavailable trades availability for speed; raising maxSurge trades cost for availability.

Caveat: externalTrafficPolicy: Local does not work with parallel node drains, so it constrains aggressive surge configurations.

For rollback procedures and maintenance windows, see the gke-upgrades skill.

Golden Path Config

See golden-path-autopilot.yaml for the full cluster-level policy settings.

來源與署名

來源:google/skills位於skills/cloud/gke-golden-path提交55b4e13

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 google/skills 的技能

Dpop Adoption

google

精選

指導為 Google OAuth 平台實作 OAuth 2.0 DPoP(RFC 9449)傳送方約束的更新權杖。

Security21K今天更新

Finding Google Skills

google

精選

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

待分類21K今天更新

Spanner Basics

google

精選

指導 Google Cloud Spanner 的執行個體與資料庫管理、結構定義設計、查詢與效能診斷。

Data & Analytics21K今天更新

Secops Triage

google

精選

引導 SOC 分析師對 Google SecOps 安全警示進行分診,從調查到結案或升級。

Security21K今天更新

Secops Investigate

google

精選

指導 SOC 分析師在 Google SecOps 中使用 UDM 查詢與時間軸進行深入的安全事件與實體調查。

Security21K今天更新

Secops Hunt

google

精選

指導在 Google SecOps 中使用 UDM 查詢、IoC 回溯、普遍性與異常分析進行主動威脅狩獵。

Security21K今天更新