GKE Batch & HPC Workloads
This reference covers running batch processing and high-performance computing (HPC) workloads on GKE.
MCP Tools:
apply_k8s_manifest,get_k8s_resource,describe_k8s_resource,get_k8s_logs,delete_k8s_resource,list_k8s_events
When to Use
- Running batch data processing pipelines
- HPC simulations (CFD, molecular dynamics, financial modeling)
- Large-scale parallel computation (MPI, MapReduce)
- ML training jobs
- CI/CD build farms
Batch Processing on GKE
Kubernetes Jobs
JobSet (for Complex Multi-Job Workflows)
The golden path enables JobSet monitoring (JOBSET in monitoringConfig).
Kueue (Job Queuing)
Kueue manages job scheduling and resource allocation for batch workloads:
HPC on GKE
Compact Placement (Low-Latency Networking)
For tightly-coupled HPC workloads that need low-latency inter-node communication:
MPI Workloads
Use the MPI Operator for MPI-based HPC applications:
Cost Optimization for Batch/HPC
Spot VMs for Batch
Batch workloads are ideal Spot VM candidates (interruptible, can checkpoint).
Use a ComputeClass with Spot-first priority and activeMigration to return to
Spot when available. See the gke-compute-classes skill for the
Spot-with-fallback pattern.
Scale-to-Zero
For batch clusters, allow node pools to scale to zero when no jobs are running:
- Autopilot (golden path): Automatic, nodes scale to zero when no pods are scheduled
- Standard: Set
--min-nodes 0on batch node pools
Best Practices & Production Guidelines
- Resource Quotas: Always specify resource requests and limits (CPU, memory, and optionally GPU/TPU) for all batch/HPC manifests. This is critical for Kueue admission, autoscaling, and preventing resource starvation in the cluster.
- TPU/Spot Cluster Maintenance: For long-running AI training runs on Spot VMs/TPUs, advise using GKE maintenance exclusions to block automatic cluster upgrades/reboots during the active training window to minimize unnecessary preemption.
- MPI Workloads: Use the Kubeflow Training Operator to orchestrate
distributed MPI applications via the
MPIJobcustom resource. - Kueue & JobSet: Use Kueue for multi-tenant job queueing and fair sharing; use JobSet for multi-component tightly coupled workloads.
- Resilience: Always set a
backoffLimiton Jobs, and implement application-level checkpointing (e.g., using Orbax or PyTorch checkpointing) to survive Spot VM preemption.

