Gke Batch Hpc

作者 google55b4e13eba6d無授權條款21K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Runs batch and HPC workloads on GKE, utilizing job queues and parallel processing. Use when running GKE batch jobs, configuring GKE HPC, or setting up GKE job queues. Don't use for standard web application deployments (use gke-app-onboarding instead).

精選僅含說明DevOps & Cloud
AI 產生的概覽

指導在 GKE 上使用 Job、JobSet、Kueue、MPI 與緊密配置執行批次與 HPC 工作負載。

功能
這份參考文件說明如何在 Google Kubernetes Engine 上執行批次處理與高效能運算工作負載。它提供 Kubernetes Job、JobSet、Kueue 佇列、緊密配置節點集區與 MPIJob 工作負載的資訊清單與指令範例,並附上 Spot VM、縮減至零與檢查點等成本與韌性建議。它產出的是設定指引,而非檔案或指令碼。
適用情境
適用於在 GKE 上執行批次資料管線、HPC 模擬、大規模平行運算、機器學習訓練工作或 CI/CD 建置叢集。它不適用於標準網頁應用程式部署。
執行需求
需要 GKE 叢集以及 kubectl、gcloud 等 Kubernetes 工具,還需要用來套用與檢視 Kubernetes 資源的 MCP 工具。安裝 Kueue 與 MPI Operator 需要連線至其資訊清單的網路存取。不隨附指令碼,僅為說明性內容。

GKE Batch & HPC Workloads

This reference covers running batch processing and high-performance computing (HPC) workloads on GKE.

MCP Tools: apply_k8s_manifest, get_k8s_resource, describe_k8s_resource, get_k8s_logs, delete_k8s_resource, list_k8s_events

When to Use

  • Running batch data processing pipelines
  • HPC simulations (CFD, molecular dynamics, financial modeling)
  • Large-scale parallel computation (MPI, MapReduce)
  • ML training jobs
  • CI/CD build farms

Batch Processing on GKE

Kubernetes Jobs

yaml
apiVersion: batch/v1kind: Jobmetadata:  name: batch-jobspec:  parallelism: 10  completions: 100  backoffLimit: 3  template:    spec:      containers:      - name: worker        image: <IMAGE>        resources:          requests:            cpu: "1"            memory: "2Gi"      restartPolicy: Never

JobSet (for Complex Multi-Job Workflows)

The golden path enables JobSet monitoring (JOBSET in monitoringConfig).

yaml
apiVersion: jobset.x-k8s.io/v1alpha2kind: JobSetmetadata:  name: training-jobspec:  replicatedJobs:  - name: workers    replicas: 4    template:      spec:        parallelism: 1        completions: 1        template:          spec:            containers:            - name: worker              image: <IMAGE>              resources:                requests:                  cpu: "4"                  memory: "8Gi"

Kueue (Job Queuing)

Kueue manages job scheduling and resource allocation for batch workloads:

bash
# Install Kueuekubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/latest/download/manifests.yaml
yaml
# Define a ClusterQueueapiVersion: kueue.x-k8s.io/v1beta1kind: ClusterQueuemetadata:  name: batch-queuespec:  namespaceSelector: {}  resourceGroups:  - coveredResources: ["cpu", "memory"]    flavors:    - name: default      resources:      - name: "cpu"        nominalQuota: 100      - name: "memory"        nominalQuota: "200Gi"---# Allow a namespace to use the queueapiVersion: kueue.x-k8s.io/v1beta1kind: LocalQueuemetadata:  name: batch-local  namespace: batch-jobsspec:  clusterQueue: batch-queue

HPC on GKE

Compact Placement (Low-Latency Networking)

For tightly-coupled HPC workloads that need low-latency inter-node communication:

bash
# Standard clusters: create node pool with compact placementgcloud container node-pools create hpc-pool \  --cluster <CLUSTER_NAME> --region <REGION> \  --machine-type c3-standard-44 \  --placement-type COMPACT \  --num-nodes 8 \  --enable-autoscaling --min-nodes 0 --max-nodes 16 \  --quiet

MPI Workloads

Use the MPI Operator for MPI-based HPC applications:

bash
# Install MPI Operatorkubectl apply -f https://raw.githubusercontent.com/kubeflow/mpi-operator/master/deploy/v2beta1/mpi-operator.yaml
yaml
apiVersion: kubeflow.org/v2beta1kind: MPIJobmetadata:  name: hpc-simulationspec:  slotsPerWorker: 4  mpiReplicaSpecs:    Launcher:      replicas: 1      template:        spec:          containers:          - name: launcher            image: <MPI_IMAGE>            command: ["mpirun", "-np", "32", "./simulation"]            resources:              requests:                cpu: "1"                memory: "2Gi"              limits:                cpu: "2"                memory: "4Gi"    Worker:      replicas: 8      template:        spec:          containers:          - name: worker            image: <MPI_IMAGE>            resources:              requests:                cpu: "4"                memory: "8Gi"              limits:                cpu: "8"                memory: "16Gi"

Cost Optimization for Batch/HPC

Spot VMs for Batch

Batch workloads are ideal Spot VM candidates (interruptible, can checkpoint). Use a ComputeClass with Spot-first priority and activeMigration to return to Spot when available. See the gke-compute-classes skill for the Spot-with-fallback pattern.

Scale-to-Zero

For batch clusters, allow node pools to scale to zero when no jobs are running:

  • Autopilot (golden path): Automatic, nodes scale to zero when no pods are scheduled
  • Standard: Set --min-nodes 0 on batch node pools

Best Practices & Production Guidelines

  • Resource Quotas: Always specify resource requests and limits (CPU, memory, and optionally GPU/TPU) for all batch/HPC manifests. This is critical for Kueue admission, autoscaling, and preventing resource starvation in the cluster.
  • TPU/Spot Cluster Maintenance: For long-running AI training runs on Spot VMs/TPUs, advise using GKE maintenance exclusions to block automatic cluster upgrades/reboots during the active training window to minimize unnecessary preemption.
  • MPI Workloads: Use the Kubeflow Training Operator to orchestrate distributed MPI applications via the MPIJob custom resource.
  • Kueue & JobSet: Use Kueue for multi-tenant job queueing and fair sharing; use JobSet for multi-component tightly coupled workloads.
  • Resilience: Always set a backoffLimit on Jobs, and implement application-level checkpointing (e.g., using Orbax or PyTorch checkpointing) to survive Spot VM preemption.

來源與署名

來源:google/skills位於skills/cloud/gke-batch-hpc提交55b4e13

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 google/skills 的技能

Dpop Adoption

google

精選

指導為 Google OAuth 平台實作 OAuth 2.0 DPoP(RFC 9449)傳送方約束的更新權杖。

Security21K今天更新

Finding Google Skills

google

精選

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

待分類21K今天更新

Spanner Basics

google

精選

指導 Google Cloud Spanner 的執行個體與資料庫管理、結構定義設計、查詢與效能診斷。

Data & Analytics21K今天更新

Secops Triage

google

精選

引導 SOC 分析師對 Google SecOps 安全警示進行分診,從調查到結案或升級。

Security21K今天更新

Secops Investigate

google

精選

指導 SOC 分析師在 Google SecOps 中使用 UDM 查詢與時間軸進行深入的安全事件與實體調查。

Security21K今天更新

Secops Hunt

google

精選

指導在 Google SecOps 中使用 UDM 查詢、IoC 回溯、普遍性與異常分析進行主動威脅狩獵。

Security21K今天更新