Gke Ai Troubleshooting Tpu Mxla Hang

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang

作者 googlec6c7e67107f02e3423df8481272f296bbe61902a无许可证21K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库今天更新

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).

精选仅含说明DevOps & Cloud
  1. c6c7e67107f02e3423df8481272f296bbe61902a当前提交 c6c7e67发布于 2026年10月9日

来源与署名

来源:google/skills位于skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang提交c6c7e67

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架