Gke Ai Troubleshooting Tpu Mxla Hang

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang

by googlec6c7e67107f02e3423df8481272f296bbe61902aNo license21K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated today

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).

FeaturedInstructions onlyDevOps & Cloud

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/path-a-compiler-hlo-divergence.md964 Btext/markdown
references/path-b-program-queueing-input-stall.md1.4 KBtext/markdown
references/path-c-hardware-network-faults.md2.5 KBtext/markdown
references/path-d-no-detected-or-unknown.md1.5 KBtext/markdown
references/path-e-healthy-telemetry.md272 Btext/markdown
SKILL.md14.5 KBtext/markdown

Source and attribution

Source:google/skillsinskills/cloud/gke-ai-troubleshooting-tpu-mxla-hangat commitc6c7e67

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal