Gke Ai Troubleshooting Tpu Performance Degradation

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation

作者 googlec6c7e67107f02e3423df8481272f296bbe61902a无许可证21K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库今天更新

Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system metrics (`kubernetes.io/node/accelerator/*`). Distinguishes hardware and network fabric throttling from workload resource bottlenecks (HBM capacity, host memory, or host CPU saturation). Use when TPU training throughput or duty cycle drops without crashing pods, when `PERFORMANCE_DEGRADATION` monitored events fire, or when triaging slow multi-slice training steps. Don't use for complete multi-slice XLA execution stalls with `HANG_DETECTED` logs (use gke-ai-troubleshooting-tpu-mxla-hang) or pod eviction/interruption restarts (use gke-ai-troubleshooting-jobset-interruption).

精选仅含说明DevOps & Cloud
  1. c6c7e67107f02e3423df8481272f296bbe61902a当前提交 c6c7e67发布于 2026年10月9日

来源与署名

来源:google/skills位于skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation提交c6c7e67

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架