Troubleshoot GKE TPU multi-slice hangs with the MXLA Hang Analyzer
Diagnose Cloud TPU multi-slice training hangs on Google Kubernetes Engine (GKE)
by correlating Megascale HANG_DETECTED logs and ML Diagnostics Workload
Monitoring Megascale XLA (MXLA) Hang Analyzer reports with 1-minute
multi-slice latency metrics (kubernetes.io/container/multislice/*) and GKE
node topology labels.
Prerequisites
- Tools: Install the
Google Cloud SDK (
gcloudwithalphacomponent forgcloud alpha mldiagnostics) andkubectl. - Cloud Billing & Project Configuration: Verify an active billing account is
linked (
gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensurecontainer.googleapis.com,logging.googleapis.com,monitoring.googleapis.com, andhypercomputecluster.googleapis.comare enabled. - Supported workloads and versions: Google Cloud ML Diagnostics only
supports JAX on TPUs (see
ML Diagnostics platform).
Workload Monitoring is enabled by default, supports the
jobsetandjobGKE job types, and is compatible with GKE versions1.36.0-gke.4681000and later (Configure GKE for ML Diagnostics, which also covers the cluster setup needed for on-demand profiling in Step 4 Path B). If a workload uses another framework (such as PyTorch) or another custom resource type,gcloud alpha mldiagnosticswon't list ML runs or monitored events for it. The Megascale XLA hang analyzer and Megascale XLA metrics require LibTPU0.40.0or later (see "Get started" in Workload monitoring with ML Diagnostics). - Required IAM Roles:
- Cluster Director Editor (
roles/hypercomputecluster.editor), the role listed in the "IAM permissions" section of ML Diagnostics platform, for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions) - Monitoring Viewer (
roles/monitoring.viewer) for the PromQL queries in Step 2 - Logs Viewer (
roles/logging.viewer) for the Cloud Logging query in Step 1 - Kubernetes Engine Viewer (
roles/container.viewer) for thekubectl get nodesquery in Step 3 - For remediation (
[High Risk]steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
- Cluster Director Editor (
- Reference Documentation:
- ML Diagnostics platform (Sections: "IAM permissions")
- Configure GKE for ML Diagnostics (Sections: "Set up with gcloud CLI, Google Cloud console, or Terraform", "Manual installation", "Connection-operator")
- Workload monitoring with ML Diagnostics (Sections: "Get started", "Megascale XLA (MXLA) Hang Analyzer", "Access Workload Monitoring information through the API", "System Metrics")
- Get started with the ML Diagnostics CLI (Sections: "List machine learning runs", "Monitored-events commands", "List profiler targets", "Capture on-demand profiler sessions")
- Get support (Sections: "Before you contact support")
- Auto-repair nodes (Sections: "Repair criteria", "Verify node auto-repair is enabled for a Standard node pool", "Enable auto-repair for an existing Standard node pool", "Node auto repair in TPU slice nodes")
- Deploy TPU workloads in GKE Standard (Sections: "Configure auto repair for TPU slice nodes")
- Dump HLO Computations (OpenXLA)
Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
MXLA Hang Analyzer routing overview
A Megascale hang occurs when a multi-slice worker has waited on a Megascale
communication operation for a set timeout period. The TPU logs then show a
Megascale HANG_DETECTED message. HANG_DETECTED is a catch-all signal that
the workload isn't progressing, and the cause can be in software or in hardware.
When a hang occurs, ML Diagnostics runs the Megascale XLA (MXLA) Hang Analyzer, which reports the likely cause as a code. Consult the
Megascale XLA (MXLA) Hang Analyzer
section for the definition and recommended action of each code, and route by
category:
- Category A: compiler or HLO divergence (never cordon or replace nodes):
- When the analyzer reports different HLO modules, inconsistent HLO
compilation, or an inconsistent launch order across VMs (such as
FINGERPRINT_MISMATCH), route to Path A: Compiler or HLO divergence [blocked].
- When the analyzer reports different HLO modules, inconsistent HLO
compilation, or an inconsistent launch order across VMs (such as
- Category B: host program queueing or data input stall (never cordon or
replace nodes):
- When the analyzer reports that VMs aren't queuing programs to the TPU or
that data input stalled (such as
DATA_INPUT_STALL), route to Path B: Host program queueing or data input stall [blocked].
- When the analyzer reports that VMs aren't queuing programs to the TPU or
that data input stalled (such as
- Category C: hardware or network faults on specific instances:
- When the analyzer attributes the hang to a TPU chip, SparseCore, ICI, or DCN networking issue, or to an unrecoverable error on specific instances, route to Path C: Hardware or network faults [blocked].
- Category D: hang signal or event exists, but the analyzer is
NOT_DETECTEDor reportsUNKNOWN:- When
HANG_DETECTEDlogs or a hang monitored event fired, but the analyzer report hasdetectionState: "NOT_DETECTED"or reportsUNKNOWN("The MXLA hang was detected but a potential cause is not determined"), route to Path D: No DETECTED analyzer or UNKNOWN cause [blocked].
- When
Diagnostic workflow
Step 0: Collect context and set the investigation window [Low Risk]
Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
- 30m]
around{issue_time}`:
{project_id}: Google Cloud project ID{location}: Google Cloud region where the ML run and GKE cluster reside (for example,us-central1){cluster_name}: GKE cluster name{namespace}/{workload_name}: Kubernetes namespace and JobSet/Pod prefix{ml_run_id}: ML Diagnostics run ID{issue_time}: Timestamp when the hang occurred (T, ISO-8601 UTC){start_time}:T - 30m{end_time}:T + 30m
Step 1: Query HANG_DETECTED logs and hang events [Low Risk]
- Check GKE container logs for
HANG_DETECTED(read-only Cloud Logging LQL): Queryk8s_containerlogs over[{start_time}, {end_time}]to confirmHANG_DETECTEDand identify the first stalled pods:
- List active or recent ML runs: Follow the
List machine learning runs
section of the ML Diagnostics CLI reference (
gcloud alpha mldiagnostics machine-learning-run list) to identify{ml_run_id}. - List and describe hang monitored events: Follow the
Monitored-events commands
section (
gcloud alpha mldiagnostics monitored-events listandgcloud alpha mldiagnostics monitored-events describe) or Access Workload Monitoring information through the API to inspect theMegascale XLA (MXLA) Hang Analyzerreport (detectionState,details, andrecommendedActions).- If
HANG_DETECTEDlogs or a hang monitored event fired, run Step 2 to correlate with the 1-minute metrics; if no analyzer reportsdetectionState: "DETECTED"or the analyzer reportsUNKNOWN, follow Path D: No DETECTED analyzer or UNKNOWN cause [blocked]. - If no
HANG_DETECTEDlog or hang monitored event exists and multi-slice latencies are normal in Step 2, rule out an MXLA hang by following Path E: Healthy telemetry [blocked].
- If
Step 2: Correlate with 1-minute multi-slice and TPU metrics [Low Risk]
Consult the
System Metrics
section of the Workload Monitoring guide for the 1-minute multi-slice network
(kubernetes.io/container/multislice/network/*), multi-slice accelerator
(kubernetes.io/container/multislice/accelerator/*), and node duty-cycle
(kubernetes.io/node/accelerator/duty_cycle) metrics, and run read-only PromQL
queries over [{start_time}, {end_time}]:
Step 3: Map culprit Compute Engine instance IDs to GKE nodes [Low Risk]
When the Megascale XLA (MXLA) Hang Analyzer reports culprit numeric Compute
Engine instance IDs in details or recommendedActions, map those numeric IDs
to GKE Node names and physical topology blocks using this read-only kubectl
query inspecting container.googleapis.com/instance_id:
Step 4: Remediation by root-cause category
Load and follow only the reference file that matches the root-cause category from Step 1:
- Path A (compiler or HLO divergence — different HLO modules, inconsistent HLO
compilation, or inconsistent launch order across VMs, such as
FINGERPRINT_MISMATCH): Read Path A: Compiler or HLO divergence [blocked]. - Path B (host program queueing or data input stall —
PROGRAM_NOT_QUEUEDorDATA_INPUT_STALL): Read Path B: Host program queueing or data input stall [blocked]. - Path C (hardware, SparseCore, ICI, DCN networking faults, or
UNRECOVERABLE_ERRORon specific instances): Read Path C: Hardware or network faults [blocked]. - Path D (
HANG_DETECTEDlogs or hang event fired, but the analyzer isNOT_DETECTEDor reportsUNKNOWN): Read Path D: No DETECTED analyzer or UNKNOWN cause [blocked]. - Path E (no
HANG_DETECTEDlogs or hang events, and steady telemetry): Read Path E: Healthy telemetry [blocked].
Guardrails
- Never change GKE-managed instance groups or VMs through Compute Engine:
Don't run
gcloud compute instance-groups managedcommands, such asdelete, on a node pool's managed instance group. Handle nodes through GKE, as described in Path C: Hardware or network faults [blocked]. - Never cordon or replace nodes for compiler divergence or input stalls: If the analyzer reports a compiler or HLO divergence, a program queueing issue, or a data input stall, don't cordon or replace TPU nodes. Follow Path A: Compiler or HLO divergence [blocked] or Path B: Host program queueing or data input stall [blocked] instead.


