Gke Workload Scaling

作者 google55b4e13eba6d無授權條款21K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Manages scaling for GKE workloads using HPA and VPA. Use when configuring Horizontal Pod Autoscaler (HPA), configuring Vertical Pod Autoscaler (VPA), or applying best practices for GKE workload autoscaling. Do not use for cluster-level autoscaling (Cluster Autoscaler), static cluster sizing, or configuring node-level machine styles directly.

精選僅含說明DevOps & Cloud
AI 產生的概覽

用於 GKE 工作負載擴縮容的工作流程與最佳實務,涵蓋手動擴縮容、HPA 與 VPA。

功能
提供在 Google Kubernetes Engine 上擴縮應用程式的逐步工作流程:手動調整副本數、依據 CPU、記憶體或自訂指標的 Horizontal Pod Autoscaler,以及用於合理配置 CPU 與記憶體的 Vertical Pod Autoscaler。內容包含指令、HPA 與 VPA 的範例 YAML 資訊清單、更新模式說明,以及附建議表的資源合理配置流程。也列出最佳實務,例如定義資源請求、避免指標衝突和使用 Pod 中斷預算。
適用情境
適用於在 GKE 上設定 Horizontal Pod Autoscaler 或 Vertical Pod Autoscaler,或為 GKE 工作負載套用自動擴縮容最佳實務。不適用於叢集層級自動擴縮容、靜態叢集規模規劃或直接設定節點機型。
執行需求
需要可存取 GKE 叢集的 kubectl 與 gcloud,HPA 需要 Metrics Server 執行,VPA 需在叢集上啟用。不包含指令碼,附帶兩個範例 YAML 資訊清單。

GKE Workload Scaling

This skill provides workflows and best practices for scaling applications on Google Kubernetes Engine (GKE). It covers manual scaling, Horizontal Pod Autoscaling (HPA), and Vertical Pod Autoscaling (VPA).

Workflows

1. Manual Scaling

Scale a deployment to a fixed number of replicas. Useful for immediate manual intervention or testing.

Command:

bash
kubectl scale deployment {deployment_name} --replicas={number} -n {namespace}
# Verify the scale eventkubectl get deployment {deployment_name} -n {namespace}

2. Horizontal Pod Autoscaling (HPA)

Automatically scale the number of pods based on observed CPU utilization, memory utilization, or custom metrics.

Prerequisites:

  • Metrics Server must be running (enabled by default on GKE).
  • Containers clearly define resource requests/limits.

Quick Command:

bash
kubectl autoscale deployment {deployment_name} --cpu-percent=50 --min=1 --max=10

Manifest Approach (Recommended): Use a YAML manifest for version-controlled configuration. See assets/hpa-example.yaml [blocked] for a template.

bash
kubectl apply -f assets/hpa-example.yaml
# Verify HPA is created and fetching metricskubectl get hpa

Custom Metrics & External Metrics: For GKE, the modern and recommended approach for scaling based on Cloud Monitoring metrics (e.g., Pub/Sub queue length) is to use the External metric type, which is natively supported by the GKE control plane without requiring the Custom Metrics Adapter. For application-specific metrics exposed via Prometheus, you can use Google Cloud Managed Service for Prometheus or the Prometheus Adapter.

3. Vertical Pod Autoscaling (VPA)

Automatically adjust the CPU and memory reservations for your pods to match actual usage. This is critical for right-sizing workloads.

Prerequisites:

  • VPA must be enabled on the cluster.
    • Autopilot: Enabled by default.
    • Standard: Must be enabled manually.

Enable VPA on Standard Cluster:

bash
gcloud container clusters update {cluster_name} --enable-vertical-pod-autoscaling --zone {zone}

Update Modes:

  • Off: Calculates recommendations but does not apply them. Good for "dry run" analysis.
  • Initial: Assigns resources only at pod creation time.
  • Auto: Updates running pods by restarting them if recommendations differ significantly from requests.
  • InPlaceOrRecreate: Attempts to update Pod resources without recreating the Pod. If in-place update is not possible, it reverts to Auto mode (requires GKE 1.34+).

Example: See assets/vpa-example.yaml [blocked] for a configuration template.

Best Practices

  1. Define Resource Requests: HPA and VPA rely on accurate resource requests. Always define them in your container specs.
  2. Avoid Metric Conflicts: Do not configure HPA and VPA to use the same metric (e.g., both CPU). This causes thrashing.
    • Typical Pattern: HPA on CPU, VPA on Memory.
  3. Pod Disruption Budgets (PDBs): Define PDBs to ensure application availability during scaling events or node upgrades.
  4. HPA Lag: HPA has a stabilization window (default 5 mins) to prevent rapid fluctuation.
  5. VPA "Auto" Mode Risks: In "Auto" mode, VPA restarts pods to change resources. Ensure your application handles restarts gracefully (e.g., handles SIGTERM).
    • Note: By default, VPA requires at least 2 replicas to perform evictions (to prevent a situation where the only running replica is evicted, causing downtime). In GKE 1.22+, you can override this by setting minReplicas in PodUpdatePolicy.

Rightsizing Workflow

  1. Deploy VPA in Off mode for 24+ hours
  2. Read recommendations: kubectl describe vpa {deployment_name}-vpa -n {namespace}
  3. Compare target values against current requests
  4. Apply with 20% buffer: new_request = target * 1.2
  5. Use patch format or update deployment manifest to apply new resource requests
ConditionRecommendationRisk
CPU request >5x P95 actualReduce to P95 * 1.2Medium
Memory request >3x P95 actualReduce to P95 * 1.2Medium
CPU request >2x P95 actualRightsizing with 20% bufferLow
No resource limits setAdd limits to prevent noisy-neighborLow

來源與署名

來源:google/skills位於skills/cloud/gke-workload-scaling提交55b4e13

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 google/skills 的技能

Dpop Adoption

google

精選

指導為 Google OAuth 平台實作 OAuth 2.0 DPoP(RFC 9449)傳送方約束的更新權杖。

Security21K今天更新

Finding Google Skills

google

精選

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

待分類21K今天更新

Spanner Basics

google

精選

指導 Google Cloud Spanner 的執行個體與資料庫管理、結構定義設計、查詢與效能診斷。

Data & Analytics21K今天更新

Secops Triage

google

精選

引導 SOC 分析師對 Google SecOps 安全警示進行分診,從調查到結案或升級。

Security21K今天更新

Secops Investigate

google

精選

指導 SOC 分析師在 Google SecOps 中使用 UDM 查詢與時間軸進行深入的安全事件與實體調查。

Security21K今天更新

Secops Hunt

google

精選

指導在 Google SecOps 中使用 UDM 查詢、IoC 回溯、普遍性與異常分析進行主動威脅狩獵。

Security21K今天更新