Gke Ai Troubleshooting Tpu Performance Degradation

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation

作者 googlec6c7e67107f02e3423df8481272f296bbe61902a無授權條款21K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫今天更新

Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system metrics (`kubernetes.io/node/accelerator/*`). Distinguishes hardware and network fabric throttling from workload resource bottlenecks (HBM capacity, host memory, or host CPU saturation). Use when TPU training throughput or duty cycle drops without crashing pods, when `PERFORMANCE_DEGRADATION` monitored events fire, or when triaging slow multi-slice training steps. Don't use for complete multi-slice XLA execution stalls with `HANG_DETECTED` logs (use gke-ai-troubleshooting-tpu-mxla-hang) or pod eviction/interruption restarts (use gke-ai-troubleshooting-jobset-interruption).

精選僅含說明DevOps & Cloud
AI 產生的概覽

使用 ML Diagnostics 事件與 Cloud Monitoring 指標診斷 GKE Cloud TPU 訓練輸送量下降。

功能
此技能引導代理執行一套唯讀診斷流程,用於排查 GKE 上 Cloud TPU 訓練輸送量下降與單步耗時回退,判定標準為 TPU 工作週期下降 15% 以上。它會將 ML Diagnostics 工作負載監控的受監控事件與分析器報告,和 1 分鐘粒度的 Cloud Monitoring 系統指標及 GKE 節點拓撲標籤相互關聯。依觸發的分析器,將排查導向四條參考路徑之一,分別對應工作負載資源瓶頸、基礎架構或網路結構節流、事件觸發但無分析器偵測到,以及遙測正常。最後產出診斷結論與供使用者自行套用的修復建議。
適用情境
適用於 TPU 訓練輸送量或工作週期下降但 Pod 未當掉、觸發 PERFORMANCE_DEGRADATION 受監控事件,或排查多切片訓練單步變慢的情況。不適用於帶有 HANG_DETECTED 記錄的完整多切片 XLA 執行停滯,也不適用於 Pod 驅逐或中斷重啟。
執行需求
需要安裝含 alpha 元件的 Google Cloud SDK(用於 gcloud alpha mldiagnostics)與 kubectl、已驗證且連結帳單專案的 gcloud 工作階段,並啟用 container、monitoring 與 hypercomputecluster API。需要 Cluster Director Editor、Monitoring Viewer、Kubernetes Engine Viewer 等 IAM 角色,修復時還需 Kubernetes Engine Cluster Admin。工作負載監控僅支援 TPU 上的 JAX,工作類型為 jobset 與 job,GKE 版本須為 1.36.0-gke.4681000 以上。不附帶指令碼,僅為說明文件。

Troubleshoot GKE TPU performance degradation with ML Diagnostics Workload Monitoring

Diagnose and mitigate Cloud TPU training throughput drops and step-time regressions (15%+ drop in TPU duty cycle) on Google Kubernetes Engine (GKE) by correlating ML Diagnostics Workload Monitoring MonitoredEvent analyzer reports with 1-minute Cloud Monitoring system metrics and GKE node topology labels.


Prerequisites

  • Tools: Install the Google Cloud SDK (gcloud with alpha component for gcloud alpha mldiagnostics) and kubectl.
  • Cloud Billing & Project Configuration: Verify an active billing account is linked (gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com, monitoring.googleapis.com, and hypercomputecluster.googleapis.com are enabled.
  • Supported workloads and GKE versions: Google Cloud ML Diagnostics only supports JAX on TPUs (see ML Diagnostics platform). Workload Monitoring is enabled by default, supports the jobset and job GKE job types, and is compatible with GKE versions 1.36.0-gke.4681000 and later, as stated in Configure GKE for ML Diagnostics. If a workload uses another framework (such as PyTorch) or another custom resource type, gcloud alpha mldiagnostics won't list ML runs or monitored events for it. On-demand profiling (Step 4 Path A) additionally requires the cluster setup described in that document.
  • Required IAM Roles:
    • Cluster Director Editor (roles/hypercomputecluster.editor), the role listed in the "IAM permissions" section of ML Diagnostics platform, for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions)
    • Monitoring Viewer (roles/monitoring.viewer) for the PromQL queries in Step 2
    • Kubernetes Engine Viewer (roles/container.viewer) for the kubectl get nodes query in Step 3
    • For remediation ([High Risk] steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
  • Reference Documentation:

Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.

When you recommend a fix, link the doc section that describes it.


Analyzer routing overview

By default, Workload Monitoring treats a 15% drop in TPU duty cycle as a performance degradation, raises a MonitoredEvent (type: PERFORMANCE_DEGRADATION), and runs its analyzers. Consult the Workload Monitoring and analyzers section in the official Cloud TPU documentation for the complete analyzer definitions, detection criteria, and subsections, and route remediation by category:

  • Category A: workload resource bottlenecks (never cordon or replace nodes):
    • When the firing analyzer (detectionState: "DETECTED") indicates a workload memory or host CPU capacity bottleneck (see the "HBM Capacity Analyzer", "Host Memory Utilization Analyzer", and "CPU Utilization Analyzer" sections of Workload monitoring with ML Diagnostics), route to Path A: Workload resource bottlenecks [blocked] (workload optimization and profiling).
  • Category B: infrastructure and network fabric throttling (culprit nodes):
    • When the firing analyzer (detectionState: "DETECTED") indicates an interconnect, thermal/power throttling, memory bandwidth, or Top-of-Rack (ToR) network fault on specific instances (see Workload Monitoring and analyzers), route to Path B: Infrastructure or network fabric throttling [blocked] (map the culprit Compute Engine instance IDs to GKE nodes, handle node repair through GKE, and escalate to Google Cloud Support).
  • Category C: PERFORMANCE_DEGRADATION event fired, but no analyzer reports detectionState: "DETECTED":
    • When a PERFORMANCE_DEGRADATION event exists (as in the sample output in List all monitored events, where analyzerReports entries have detectionState: "NOT_DETECTED"), route to Path C: Event fired with no DETECTED analyzer [blocked].

Diagnostic workflow

Step 0: Collect context and set the investigation window [Low Risk]

Collect the target parameters. By default, query a 60-minute window `[T - 30m, T

  • 30m]around{issue_time}`:
  • {project_id}: Google Cloud project ID
  • {location}: Google Cloud region where the ML run and GKE cluster reside (for example, us-central1)
  • {cluster_name}: GKE cluster name
  • {workload_name} / {ml_run_id}: JobSet / workload name or ML Diagnostics run ID
  • {issue_time}: Timestamp when throughput degradation was observed (T, ISO-8601 UTC)
  • {start_time}: T - 30m
  • {end_time}: T + 30m

Step 1: Query ML runs and monitoredEvents [Low Risk]

  1. List active or recent ML runs: Give the user gcloud alpha mldiagnostics machine-learning-run list with a link to List machine learning runs in the ML Diagnostics CLI reference, and locate {ml_run_id} matching {workload_name}.
  2. List performance degradation events: Give the user gcloud alpha mldiagnostics monitored-events list with a link to Monitored-events commands, or the hypercomputecluster.googleapis.com/v1alpha API with a link to Access Workload Monitoring information through the API, to check for PERFORMANCE_DEGRADATION events during [{start_time}, {end_time}].
  3. Describe the MonitoredEvent: Give the user gcloud alpha mldiagnostics monitored-events describe with a link to the same Monitored-events commands section, and inspect the analyzerReports array (analyzer, detectionState, details, and recommendedActions).
    • If a PERFORMANCE_DEGRADATION event fired, run Step 2 to corroborate with the 1-minute system metrics; if none of its analyzerReports entries has detectionState: "DETECTED", follow Path C: Event fired with no DETECTED analyzer [blocked].
    • If no PERFORMANCE_DEGRADATION event exists and duty cycle is steady in Step 2, rule out TPU performance degradation by following Path D: Healthy telemetry [blocked].

Step 2: Correlate with 1-minute Cloud Monitoring system metrics [Low Risk]

Consult the System Metrics section in the Workload Monitoring documentation for the 1-minute Cloud Monitoring metrics exported for TPU and host devices, and run read-only PromQL queries over [{start_time}, {end_time}] to corroborate the analyzer report:

promql
# 1. Node TPU duty cycle (look for the drop on the affected nodes)kubernetes_io:node_accelerator_duty_cycle{  monitored_resource="k8s_node",  project_id="{project_id}",  cluster_name="{cluster_name}"}
# 2. HBM utilization ratio by node (around 0.90 is approaching the limit)sum by (node_name) (  kubernetes_io:node_accelerator_memory_used{    monitored_resource="k8s_node",    project_id="{project_id}",    cluster_name="{cluster_name}"  })/sum by (node_name) (  kubernetes_io:node_accelerator_memory_total{    monitored_resource="k8s_node",    project_id="{project_id}",    cluster_name="{cluster_name}"  })
# 3. Host memory and CPU allocatable utilizationkubernetes_io:node_memory_allocatable_utilization{  monitored_resource="k8s_node",  project_id="{project_id}",  cluster_name="{cluster_name}"}
kubernetes_io:node_cpu_allocatable_utilization{  monitored_resource="k8s_node",  project_id="{project_id}",  cluster_name="{cluster_name}"}

Step 3: Map culprit instance IDs to GKE nodes and topology [Low Risk]

When an infrastructure analyzer reports culprit numeric Compute Engine instance IDs in details or recommendedActions, map those numeric instance IDs to GKE Node names and physical topology blocks using this read-only kubectl query inspecting container.googleapis.com/instance_id:

bash
kubectl get nodes -l cloud.google.com/gke-tpu-accelerator \  -o jsonpath='{range .items[*]}{.metadata.name}{"\tinstance_id="}{.metadata.annotations.container\.googleapis\.com/instance_id}{"\tblock="}{.metadata.labels.cloud\.google\.com/gce-topology-block}{"\tsubblock="}{.metadata.labels.cloud\.google\.com/gce-topology-subblock}{"\thost="}{.metadata.labels.cloud\.google\.com/gce-topology-host}{"\n"}{end}'

Step 4: Remediation by analyzer category

Load and follow only the reference file that matches the analyzer category from Step 1:

  • Path A (workload resource bottlenecks — HBM, host memory, or host CPU utilization): Read Path A: Workload resource bottlenecks [blocked].
  • Path B (infrastructure, thermal, ICI, memory bandwidth, or network fabric throttling): Read Path B: Infrastructure or network fabric throttling [blocked].
  • Path C (PERFORMANCE_DEGRADATION event fired with all analyzers NOT_DETECTED): Read Path C: Event fired with no DETECTED analyzer [blocked].
  • Path D (no PERFORMANCE_DEGRADATION events and steady telemetry): Read Path D: Healthy telemetry [blocked].

Guardrails

  1. Never change GKE-managed instance groups or VMs through Compute Engine: Don't run gcloud compute instance-groups managed commands, such as delete, on a node pool's managed instance group. Handle nodes through GKE, as described in Path B: Infrastructure or network fabric throttling [blocked].
  2. Never cordon nodes for workload resource saturation: If only the HBM capacity, host memory, or host CPU utilization analyzers detected an issue, don't cordon or replace nodes.

來源與署名

來源:google/skills位於skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation提交c6c7e67

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架