Gke Ai Troubleshooting Tpu Mxla Hang

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang

作者 googlec6c7e67107f02e3423df8481272f296bbe61902a无许可证21K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库今天更新

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).

精选仅含说明DevOps & Cloud

仅公开文件列表。将技能安装到工作区后即可查看文件内容。

路径大小类型
references/path-a-compiler-hlo-divergence.md964 Btext/markdown
references/path-b-program-queueing-input-stall.md1.4 KBtext/markdown
references/path-c-hardware-network-faults.md2.5 KBtext/markdown
references/path-d-no-detected-or-unknown.md1.5 KBtext/markdown
references/path-e-healthy-telemetry.md272 Btext/markdown
SKILL.md14.5 KBtext/markdown

来源与署名

来源:google/skills位于skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang提交c6c7e67

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架