Gke Ai Troubleshooting Tpu Mxla Hang

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang

by googlec6c7e67107f02e3423df8481272f296bbe61902aNo license21K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated today

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).

FeaturedInstructions onlyDevOps & Cloud
  1. c6c7e67107f02e3423df8481272f296bbe61902aCurrentcommit c6c7e67Published Oct 9, 2026

Source and attribution

Source:google/skillsinskills/cloud/gke-ai-troubleshooting-tpu-mxla-hangat commitc6c7e67

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal