Gke Workload Scaling

by google55b4e13eba6dNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Manages scaling for GKE workloads using HPA and VPA. Use when configuring Horizontal Pod Autoscaler (HPA), configuring Vertical Pod Autoscaler (VPA), or applying best practices for GKE workload autoscaling. Do not use for cluster-level autoscaling (Cluster Autoscaler), static cluster sizing, or configuring node-level machine styles directly.

FeaturedInstructions onlyDevOps & Cloud
AI-generated overview

Workflows and best practices for scaling GKE workloads with manual scaling, HPA and VPA.

What it does
Provides step-by-step workflows for scaling applications on Google Kubernetes Engine: manual replica scaling, Horizontal Pod Autoscaling based on CPU, memory or custom metrics, and Vertical Pod Autoscaling for right-sizing CPU and memory. It includes commands, example YAML manifests for HPA and VPA, update-mode guidance, and a rightsizing workflow with a recommendation table. It also lists best practices such as defining resource requests, avoiding metric conflicts, and using Pod Disruption Budgets.
When to use it
Use when configuring Horizontal Pod Autoscaler or Vertical Pod Autoscaler on GKE, or when applying autoscaling best practices to GKE workloads. Not intended for cluster-level autoscaling, static cluster sizing, or configuring node machine types directly.
Requirements
Requires kubectl and gcloud access to a GKE cluster, with Metrics Server running for HPA and VPA enabled on the cluster. Ships no scripts; includes two example YAML manifests.

GKE Workload Scaling

This skill provides workflows and best practices for scaling applications on Google Kubernetes Engine (GKE). It covers manual scaling, Horizontal Pod Autoscaling (HPA), and Vertical Pod Autoscaling (VPA).

Workflows

1. Manual Scaling

Scale a deployment to a fixed number of replicas. Useful for immediate manual intervention or testing.

Command:

bash
kubectl scale deployment {deployment_name} --replicas={number} -n {namespace}
# Verify the scale eventkubectl get deployment {deployment_name} -n {namespace}

2. Horizontal Pod Autoscaling (HPA)

Automatically scale the number of pods based on observed CPU utilization, memory utilization, or custom metrics.

Prerequisites:

  • Metrics Server must be running (enabled by default on GKE).
  • Containers clearly define resource requests/limits.

Quick Command:

bash
kubectl autoscale deployment {deployment_name} --cpu-percent=50 --min=1 --max=10

Manifest Approach (Recommended): Use a YAML manifest for version-controlled configuration. See assets/hpa-example.yaml [blocked] for a template.

bash
kubectl apply -f assets/hpa-example.yaml
# Verify HPA is created and fetching metricskubectl get hpa

Custom Metrics & External Metrics: For GKE, the modern and recommended approach for scaling based on Cloud Monitoring metrics (e.g., Pub/Sub queue length) is to use the External metric type, which is natively supported by the GKE control plane without requiring the Custom Metrics Adapter. For application-specific metrics exposed via Prometheus, you can use Google Cloud Managed Service for Prometheus or the Prometheus Adapter.

3. Vertical Pod Autoscaling (VPA)

Automatically adjust the CPU and memory reservations for your pods to match actual usage. This is critical for right-sizing workloads.

Prerequisites:

  • VPA must be enabled on the cluster.
    • Autopilot: Enabled by default.
    • Standard: Must be enabled manually.

Enable VPA on Standard Cluster:

bash
gcloud container clusters update {cluster_name} --enable-vertical-pod-autoscaling --zone {zone}

Update Modes:

  • Off: Calculates recommendations but does not apply them. Good for "dry run" analysis.
  • Initial: Assigns resources only at pod creation time.
  • Auto: Updates running pods by restarting them if recommendations differ significantly from requests.
  • InPlaceOrRecreate: Attempts to update Pod resources without recreating the Pod. If in-place update is not possible, it reverts to Auto mode (requires GKE 1.34+).

Example: See assets/vpa-example.yaml [blocked] for a configuration template.

Best Practices

  1. Define Resource Requests: HPA and VPA rely on accurate resource requests. Always define them in your container specs.
  2. Avoid Metric Conflicts: Do not configure HPA and VPA to use the same metric (e.g., both CPU). This causes thrashing.
    • Typical Pattern: HPA on CPU, VPA on Memory.
  3. Pod Disruption Budgets (PDBs): Define PDBs to ensure application availability during scaling events or node upgrades.
  4. HPA Lag: HPA has a stabilization window (default 5 mins) to prevent rapid fluctuation.
  5. VPA "Auto" Mode Risks: In "Auto" mode, VPA restarts pods to change resources. Ensure your application handles restarts gracefully (e.g., handles SIGTERM).
    • Note: By default, VPA requires at least 2 replicas to perform evictions (to prevent a situation where the only running replica is evicted, causing downtime). In GKE 1.22+, you can override this by setting minReplicas in PodUpdatePolicy.

Rightsizing Workflow

  1. Deploy VPA in Off mode for 24+ hours
  2. Read recommendations: kubectl describe vpa {deployment_name}-vpa -n {namespace}
  3. Compare target values against current requests
  4. Apply with 20% buffer: new_request = target * 1.2
  5. Use patch format or update deployment manifest to apply new resource requests
ConditionRecommendationRisk
CPU request >5x P95 actualReduce to P95 * 1.2Medium
Memory request >3x P95 actualReduce to P95 * 1.2Medium
CPU request >2x P95 actualRightsizing with 20% bufferLow
No resource limits setAdd limits to prevent noisy-neighborLow

Source and attribution

Source:google/skillsinskills/cloud/gke-workload-scalingat commit55b4e13

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from google/skills

Dpop Adoption

google

Featured

Guides implementation of OAuth 2.0 DPoP (RFC 9449) sender-constrained refresh tokens for Google's OAuth platform.

SecurityOct 8, 2026

Finding Google Skills

google

Featured

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

Awaiting classificationOct 8, 2026

Spanner Basics

google

Featured

Guides Google Cloud Spanner administration, schema design, querying and performance diagnosis.

Data & AnalyticsOct 8, 2026

Secops Triage

google

Featured

Guides SOC analysts through triaging Google SecOps security alerts, from investigation to closure or escalation.

SecurityOct 8, 2026

Secops Investigate

google

Featured

Guides SOC analysts through deep security incident and entity investigations in Google SecOps using UDM queries and timelines.

SecurityOct 8, 2026

Secops Hunt

google

Featured

Guides proactive threat hunting in Google SecOps using UDM queries, IoC lookback, prevalence and outlier analysis.

SecurityOct 8, 2026