Gke Ai Troubleshooting Tpu Mxla Hang

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang

作者 googlec6c7e67107f02e3423df8481272f296bbe61902a無授權條款21K 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫今天更新

Diagnose GKE Cloud TPU multi-slice training hangs (Megascale `HANG_DETECTED` logs and hang monitored events) using the ML Diagnostics `Megascale XLA (MXLA) Hang Analyzer` (`gcloud alpha mldiagnostics monitored-events`) and 1-minute Cloud Monitoring multi-slice latency metrics (`kubernetes.io/container/multislice/*`). Distinguishes XLA compiler/HLO launch divergence and host data-input stalls from TPU chip, SparseCore, ICI, or network fabric faults. Use when multi-slice TPU training jobs freeze without progressing steps, emit `HANG_DETECTED`, or stall in collective operations. Don't use for gradual step-time throughput drops without hangs (use gke-ai-troubleshooting-tpu-performance-degradation) or pod preemption/eviction restarts (use gke-ai-troubleshooting-jobset-interruption).

精選僅含說明DevOps & Cloud
  1. c6c7e67107f02e3423df8481272f296bbe61902a目前提交 c6c7e67發布於 2026年10月9日

來源與署名

來源:google/skills位於skills/cloud/gke-ai-troubleshooting-tpu-mxla-hang提交c6c7e67

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架