Gke Batch Hpc

作者 google55b4e13eba6d无许可证21K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job queues. Don't use for standard web application deployments (use gke-app-onboarding instead).

精选仅含说明DevOps & Cloud
AI 生成的概览

指导在 GKE 上使用 Job、JobSet、Kueue、MPI 和紧凑放置运行批处理与 HPC 工作负载。

功能
该参考文档说明如何在 Google Kubernetes Engine 上运行批处理和高性能计算工作负载。它给出 Kubernetes Job、JobSet、Kueue 队列、紧凑放置节点池和 MPIJob 工作负载的清单与命令示例,并提供 Spot VM、缩容到零和检查点等成本与韧性建议。它产出的是配置指导,而不是文件或脚本。
适用场景
适用于在 GKE 上运行批处理数据管道、HPC 仿真、大规模并行计算、机器学习训练任务或 CI/CD 构建集群。它不适用于标准 Web 应用部署。
运行要求
需要 GKE 集群以及 kubectl、gcloud 等 Kubernetes 工具,还需要用于应用和查看 Kubernetes 资源的 MCP 工具。安装 Kueue 和 MPI Operator 需要访问其清单的网络连接。不附带脚本,仅为说明性内容。

GKE Batch & HPC Workloads

This reference covers running batch processing and high-performance computing (HPC) workloads on GKE.

MCP Tools: apply_k8s_manifest, get_k8s_resource, describe_k8s_resource, get_k8s_logs, delete_k8s_resource, list_k8s_events

When to Use

  • Running batch data processing pipelines
  • HPC simulations (CFD, molecular dynamics, financial modeling)
  • Large-scale parallel computation (MPI, MapReduce)
  • ML training jobs
  • CI/CD build farms

Batch Processing on GKE

Kubernetes Jobs

yaml
apiVersion: batch/v1kind: Jobmetadata:  name: batch-jobspec:  parallelism: 10  completions: 100  backoffLimit: 3  template:    spec:      containers:      - name: worker        image: <IMAGE>        resources:          requests:            cpu: "1"            memory: "2Gi"      restartPolicy: Never

JobSet (for Complex Multi-Job Workflows)

The golden path enables JobSet monitoring (JOBSET in monitoringConfig).

yaml
apiVersion: jobset.x-k8s.io/v1alpha2kind: JobSetmetadata:  name: training-jobspec:  replicatedJobs:  - name: workers    replicas: 4    template:      spec:        parallelism: 1        completions: 1        template:          spec:            containers:            - name: worker              image: <IMAGE>              resources:                requests:                  cpu: "4"                  memory: "8Gi"

Kueue (Job Queuing)

Kueue manages job scheduling and resource allocation for batch workloads:

bash
# Install Kueuekubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/latest/download/manifests.yaml
yaml
# Define a ClusterQueueapiVersion: kueue.x-k8s.io/v1beta1kind: ClusterQueuemetadata:  name: batch-queuespec:  namespaceSelector: {}  resourceGroups:  - coveredResources: ["cpu", "memory"]    flavors:    - name: default      resources:      - name: "cpu"        nominalQuota: 100      - name: "memory"        nominalQuota: "200Gi"---# Allow a namespace to use the queueapiVersion: kueue.x-k8s.io/v1beta1kind: LocalQueuemetadata:  name: batch-local  namespace: batch-jobsspec:  clusterQueue: batch-queue

HPC on GKE

Compact Placement (Low-Latency Networking)

For tightly-coupled HPC workloads that need low-latency inter-node communication:

bash
# Standard clusters: create node pool with compact placementgcloud container node-pools create hpc-pool \  --cluster <CLUSTER_NAME> --region <REGION> \  --machine-type c3-standard-44 \  --placement-type COMPACT \  --num-nodes 8 \  --enable-autoscaling --min-nodes 0 --max-nodes 16 \  --quiet

MPI Workloads

Use the MPI Operator for MPI-based HPC applications:

bash
# Install MPI Operatorkubectl apply -f https://raw.githubusercontent.com/kubeflow/mpi-operator/master/deploy/v2beta1/mpi-operator.yaml
yaml
apiVersion: kubeflow.org/v2beta1kind: MPIJobmetadata:  name: hpc-simulationspec:  slotsPerWorker: 4  mpiReplicaSpecs:    Launcher:      replicas: 1      template:        spec:          containers:          - name: launcher            image: <MPI_IMAGE>            command: ["mpirun", "-np", "32", "./simulation"]            resources:              requests:                cpu: "1"                memory: "2Gi"              limits:                cpu: "2"                memory: "4Gi"    Worker:      replicas: 8      template:        spec:          containers:          - name: worker            image: <MPI_IMAGE>            resources:              requests:                cpu: "4"                memory: "8Gi"              limits:                cpu: "8"                memory: "16Gi"

Cost Optimization for Batch/HPC

Spot VMs for Batch

Batch workloads are ideal Spot VM candidates (interruptible, can checkpoint). Use a ComputeClass with Spot-first priority and activeMigration to return to Spot when available. See the gke-compute-classes skill for the Spot-with-fallback pattern.

Scale-to-Zero

For batch clusters, allow node pools to scale to zero when no jobs are running:

  • Autopilot (golden path): Automatic, nodes scale to zero when no pods are scheduled
  • Standard: Set --min-nodes 0 on batch node pools

Best Practices & Production Guidelines

  • Resource Quotas: Always specify resource requests and limits (CPU, memory, and optionally GPU/TPU) for all batch/HPC manifests. This is critical for Kueue admission, autoscaling, and preventing resource starvation in the cluster.
  • TPU/Spot Cluster Maintenance: For long-running AI training runs on Spot VMs/TPUs, advise using GKE maintenance exclusions to block automatic cluster upgrades/reboots during the active training window to minimize unnecessary preemption.
  • MPI Workloads: Use the Kubeflow Training Operator to orchestrate distributed MPI applications via the MPIJob custom resource.
  • Kueue & JobSet: Use Kueue for multi-tenant job queueing and fair sharing; use JobSet for multi-component tightly coupled workloads.
  • Resilience: Always set a backoffLimit on Jobs, and implement application-level checkpointing (e.g., using Orbax or PyTorch checkpointing) to survive Spot VM preemption.

来源与署名

来源:google/skills位于skills/cloud/gke-batch-hpc提交55b4e13

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 google/skills 的技能