Gke Ai Troubleshooting Tpu Performance Degradation

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation

作者 googlec6c7e67107f02e3423df8481272f296bbe61902a無授權條款21K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫今天更新

Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system metrics (`kubernetes.io/node/accelerator/*`). Distinguishes hardware and network fabric throttling from workload resource bottlenecks (HBM capacity, host memory, or host CPU saturation). Use when TPU training throughput or duty cycle drops without crashing pods, when `PERFORMANCE_DEGRADATION` monitored events fire, or when triaging slow multi-slice training steps. Don't use for complete multi-slice XLA execution stalls with `HANG_DETECTED` logs (use gke-ai-troubleshooting-tpu-mxla-hang) or pod eviction/interruption restarts (use gke-ai-troubleshooting-jobset-interruption).

精選僅含說明DevOps & Cloud
  1. c6c7e67107f02e3423df8481272f296bbe61902a目前提交 c6c7e67發布於 2026年10月9日

來源與署名

來源:google/skills位於skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation提交c6c7e67

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架