Gke Batch Hpc

by google55b4e13eba6dNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job queues. Don't use for standard web application deployments (use gke-app-onboarding instead).

FeaturedInstructions onlyDevOps & Cloud
AI-generated overview

Guides running batch and HPC workloads on GKE using Jobs, JobSet, Kueue, MPI, and compact placement.

What it does
This reference explains how to run batch processing and high-performance computing workloads on Google Kubernetes Engine. It provides manifest and command examples for Kubernetes Jobs, JobSet, Kueue queues, compact placement node pools, and MPIJob workloads, plus cost and resilience guidance such as Spot VMs, scale-to-zero, and checkpointing. It produces configuration guidance rather than files or scripts.
When to use it
Use it when running batch data pipelines, HPC simulations, large-scale parallel computation, ML training jobs, or CI/CD build farms on GKE. It is not intended for standard web application deployments.
Requirements
Requires a GKE cluster and Kubernetes tooling such as kubectl and gcloud, plus MCP tools for applying and inspecting Kubernetes resources. Installing Kueue and the MPI Operator requires network access to their manifests. No scripts are shipped; it is instructions only.

GKE Batch & HPC Workloads

This reference covers running batch processing and high-performance computing (HPC) workloads on GKE.

MCP Tools: apply_k8s_manifest, get_k8s_resource, describe_k8s_resource, get_k8s_logs, delete_k8s_resource, list_k8s_events

When to Use

  • Running batch data processing pipelines
  • HPC simulations (CFD, molecular dynamics, financial modeling)
  • Large-scale parallel computation (MPI, MapReduce)
  • ML training jobs
  • CI/CD build farms

Batch Processing on GKE

Kubernetes Jobs

yaml
apiVersion: batch/v1kind: Jobmetadata:  name: batch-jobspec:  parallelism: 10  completions: 100  backoffLimit: 3  template:    spec:      containers:      - name: worker        image: <IMAGE>        resources:          requests:            cpu: "1"            memory: "2Gi"      restartPolicy: Never

JobSet (for Complex Multi-Job Workflows)

The golden path enables JobSet monitoring (JOBSET in monitoringConfig).

yaml
apiVersion: jobset.x-k8s.io/v1alpha2kind: JobSetmetadata:  name: training-jobspec:  replicatedJobs:  - name: workers    replicas: 4    template:      spec:        parallelism: 1        completions: 1        template:          spec:            containers:            - name: worker              image: <IMAGE>              resources:                requests:                  cpu: "4"                  memory: "8Gi"

Kueue (Job Queuing)

Kueue manages job scheduling and resource allocation for batch workloads:

bash
# Install Kueuekubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/latest/download/manifests.yaml
yaml
# Define a ClusterQueueapiVersion: kueue.x-k8s.io/v1beta1kind: ClusterQueuemetadata:  name: batch-queuespec:  namespaceSelector: {}  resourceGroups:  - coveredResources: ["cpu", "memory"]    flavors:    - name: default      resources:      - name: "cpu"        nominalQuota: 100      - name: "memory"        nominalQuota: "200Gi"---# Allow a namespace to use the queueapiVersion: kueue.x-k8s.io/v1beta1kind: LocalQueuemetadata:  name: batch-local  namespace: batch-jobsspec:  clusterQueue: batch-queue

HPC on GKE

Compact Placement (Low-Latency Networking)

For tightly-coupled HPC workloads that need low-latency inter-node communication:

bash
# Standard clusters: create node pool with compact placementgcloud container node-pools create hpc-pool \  --cluster <CLUSTER_NAME> --region <REGION> \  --machine-type c3-standard-44 \  --placement-type COMPACT \  --num-nodes 8 \  --enable-autoscaling --min-nodes 0 --max-nodes 16 \  --quiet

MPI Workloads

Use the MPI Operator for MPI-based HPC applications:

bash
# Install MPI Operatorkubectl apply -f https://raw.githubusercontent.com/kubeflow/mpi-operator/master/deploy/v2beta1/mpi-operator.yaml
yaml
apiVersion: kubeflow.org/v2beta1kind: MPIJobmetadata:  name: hpc-simulationspec:  slotsPerWorker: 4  mpiReplicaSpecs:    Launcher:      replicas: 1      template:        spec:          containers:          - name: launcher            image: <MPI_IMAGE>            command: ["mpirun", "-np", "32", "./simulation"]            resources:              requests:                cpu: "1"                memory: "2Gi"              limits:                cpu: "2"                memory: "4Gi"    Worker:      replicas: 8      template:        spec:          containers:          - name: worker            image: <MPI_IMAGE>            resources:              requests:                cpu: "4"                memory: "8Gi"              limits:                cpu: "8"                memory: "16Gi"

Cost Optimization for Batch/HPC

Spot VMs for Batch

Batch workloads are ideal Spot VM candidates (interruptible, can checkpoint). Use a ComputeClass with Spot-first priority and activeMigration to return to Spot when available. See the gke-compute-classes skill for the Spot-with-fallback pattern.

Scale-to-Zero

For batch clusters, allow node pools to scale to zero when no jobs are running:

  • Autopilot (golden path): Automatic, nodes scale to zero when no pods are scheduled
  • Standard: Set --min-nodes 0 on batch node pools

Best Practices & Production Guidelines

  • Resource Quotas: Always specify resource requests and limits (CPU, memory, and optionally GPU/TPU) for all batch/HPC manifests. This is critical for Kueue admission, autoscaling, and preventing resource starvation in the cluster.
  • TPU/Spot Cluster Maintenance: For long-running AI training runs on Spot VMs/TPUs, advise using GKE maintenance exclusions to block automatic cluster upgrades/reboots during the active training window to minimize unnecessary preemption.
  • MPI Workloads: Use the Kubeflow Training Operator to orchestrate distributed MPI applications via the MPIJob custom resource.
  • Kueue & JobSet: Use Kueue for multi-tenant job queueing and fair sharing; use JobSet for multi-component tightly coupled workloads.
  • Resilience: Always set a backoffLimit on Jobs, and implement application-level checkpointing (e.g., using Orbax or PyTorch checkpointing) to survive Spot VM preemption.

Source and attribution

Source:google/skillsinskills/cloud/gke-batch-hpcat commit55b4e13

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from google/skills

Dpop Adoption

google

Featured

Guides implementation of OAuth 2.0 DPoP (RFC 9449) sender-constrained refresh tokens for Google's OAuth platform.

SecurityOct 8, 2026

Finding Google Skills

google

Featured

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

Awaiting classificationOct 8, 2026

Spanner Basics

google

Featured

Guides Google Cloud Spanner administration, schema design, querying and performance diagnosis.

Data & AnalyticsOct 8, 2026

Secops Triage

google

Featured

Guides SOC analysts through triaging Google SecOps security alerts, from investigation to closure or escalation.

SecurityOct 8, 2026

Secops Investigate

google

Featured

Guides SOC analysts through deep security incident and entity investigations in Google SecOps using UDM queries and timelines.

SecurityOct 8, 2026

Secops Hunt

google

Featured

Guides proactive threat hunting in Google SecOps using UDM queries, IoC lookback, prevalence and outlier analysis.

SecurityOct 8, 2026