Troubleshoot GKE TPU performance degradation with ML Diagnostics Workload Monitoring
Diagnose and mitigate Cloud TPU training throughput drops and step-time
regressions (15%+ drop in TPU duty cycle) on Google Kubernetes Engine (GKE) by
correlating ML Diagnostics Workload Monitoring MonitoredEvent analyzer
reports with 1-minute Cloud Monitoring system metrics and GKE node topology
labels.
Prerequisites
- Tools: Install the
Google Cloud SDK (
gcloudwithalphacomponent forgcloud alpha mldiagnostics) andkubectl. - Cloud Billing & Project Configuration: Verify an active billing account is
linked (
gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensurecontainer.googleapis.com,monitoring.googleapis.com, andhypercomputecluster.googleapis.comare enabled. - Supported workloads and GKE versions: Google Cloud ML Diagnostics only
supports JAX on TPUs (see
ML Diagnostics platform).
Workload Monitoring is enabled by default, supports the
jobsetandjobGKE job types, and is compatible with GKE versions1.36.0-gke.4681000and later, as stated in Configure GKE for ML Diagnostics. If a workload uses another framework (such as PyTorch) or another custom resource type,gcloud alpha mldiagnosticswon't list ML runs or monitored events for it. On-demand profiling (Step 4 Path A) additionally requires the cluster setup described in that document. - Required IAM Roles:
- Cluster Director Editor (
roles/hypercomputecluster.editor), the role listed in the "IAM permissions" section of ML Diagnostics platform, for the ML Diagnostics CLI and API calls in this skill (ML runs, monitored events, and on-demand profiler sessions) - Monitoring Viewer (
roles/monitoring.viewer) for the PromQL queries in Step 2 - Kubernetes Engine Viewer (
roles/container.viewer) for thekubectl get nodesquery in Step 3 - For remediation (
[High Risk]steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
- Cluster Director Editor (
- Reference Documentation:
- ML Diagnostics platform (Sections: "IAM permissions")
- Configure GKE for ML Diagnostics (Sections: "Set up with gcloud CLI, Google Cloud console, or Terraform", "Manual installation", "Connection-operator")
- Workload monitoring with ML Diagnostics (Sections: "Workload Monitoring and analyzers", "HBM Capacity Analyzer", "Host Memory Utilization Analyzer", "CPU Utilization Analyzer", "Access Workload Monitoring information through the API", "List all monitored events", "System Metrics")
- Get started with the ML Diagnostics CLI (Sections: "List machine learning runs", "Monitored-events commands", "List profiler targets", "Capture on-demand profiler sessions")
- Scale container resource requests and limits (Sections: "Identify workloads without resource requests or limits")
- Auto-repair nodes (Sections: "Repair criteria", "Verify node auto-repair is enabled for a Standard node pool", "Enable auto-repair for an existing Standard node pool", "Node auto repair in TPU slice nodes")
- Deploy TPU workloads in GKE Standard (Sections: "Configure auto repair for TPU slice nodes")
- Get support (Sections: "Before you contact support")
Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.
When you recommend a fix, link the doc section that describes it.
Analyzer routing overview
By default, Workload Monitoring treats a 15% drop in TPU duty cycle as a
performance degradation, raises a MonitoredEvent (type: PERFORMANCE_DEGRADATION), and runs its analyzers. Consult the
Workload Monitoring and analyzers
section in the official Cloud TPU documentation for the complete analyzer
definitions, detection criteria, and subsections, and route remediation by
category:
- Category A: workload resource bottlenecks (never cordon or replace nodes):
- When the firing analyzer (
detectionState: "DETECTED") indicates a workload memory or host CPU capacity bottleneck (see the "HBM Capacity Analyzer", "Host Memory Utilization Analyzer", and "CPU Utilization Analyzer" sections of Workload monitoring with ML Diagnostics), route to Path A: Workload resource bottlenecks [blocked] (workload optimization and profiling).
- When the firing analyzer (
- Category B: infrastructure and network fabric throttling (culprit nodes):
- When the firing analyzer (
detectionState: "DETECTED") indicates an interconnect, thermal/power throttling, memory bandwidth, or Top-of-Rack (ToR) network fault on specific instances (see Workload Monitoring and analyzers), route to Path B: Infrastructure or network fabric throttling [blocked] (map the culprit Compute Engine instance IDs to GKE nodes, handle node repair through GKE, and escalate to Google Cloud Support).
- When the firing analyzer (
- Category C:
PERFORMANCE_DEGRADATIONevent fired, but no analyzer reportsdetectionState: "DETECTED":- When a
PERFORMANCE_DEGRADATIONevent exists (as in the sample output in List all monitored events, whereanalyzerReportsentries havedetectionState: "NOT_DETECTED"), route to Path C: Event fired with no DETECTED analyzer [blocked].
- When a
Diagnostic workflow
Step 0: Collect context and set the investigation window [Low Risk]
Collect the target parameters. By default, query a 60-minute window `[T - 30m, T
- 30m]
around{issue_time}`:
{project_id}: Google Cloud project ID{location}: Google Cloud region where the ML run and GKE cluster reside (for example,us-central1){cluster_name}: GKE cluster name{workload_name}/{ml_run_id}: JobSet / workload name or ML Diagnostics run ID{issue_time}: Timestamp when throughput degradation was observed (T, ISO-8601 UTC){start_time}:T - 30m{end_time}:T + 30m
Step 1: Query ML runs and monitoredEvents [Low Risk]
- List active or recent ML runs: Give the user
gcloud alpha mldiagnostics machine-learning-run listwith a link to List machine learning runs in the ML Diagnostics CLI reference, and locate{ml_run_id}matching{workload_name}. - List performance degradation events: Give the user
gcloud alpha mldiagnostics monitored-events listwith a link to Monitored-events commands, or thehypercomputecluster.googleapis.com/v1alphaAPI with a link to Access Workload Monitoring information through the API, to check forPERFORMANCE_DEGRADATIONevents during[{start_time}, {end_time}]. - Describe the
MonitoredEvent: Give the usergcloud alpha mldiagnostics monitored-events describewith a link to the same Monitored-events commands section, and inspect theanalyzerReportsarray (analyzer,detectionState,details, andrecommendedActions).- If a
PERFORMANCE_DEGRADATIONevent fired, run Step 2 to corroborate with the 1-minute system metrics; if none of itsanalyzerReportsentries hasdetectionState: "DETECTED", follow Path C: Event fired with no DETECTED analyzer [blocked]. - If no
PERFORMANCE_DEGRADATIONevent exists and duty cycle is steady in Step 2, rule out TPU performance degradation by following Path D: Healthy telemetry [blocked].
- If a
Step 2: Correlate with 1-minute Cloud Monitoring system metrics [Low Risk]
Consult the
System Metrics
section in the Workload Monitoring documentation for the 1-minute Cloud
Monitoring metrics exported for TPU and host devices, and run read-only PromQL
queries over [{start_time}, {end_time}] to corroborate the analyzer report:
Step 3: Map culprit instance IDs to GKE nodes and topology [Low Risk]
When an infrastructure analyzer reports culprit numeric Compute Engine instance
IDs in details or recommendedActions, map those numeric instance IDs to GKE
Node names and physical topology blocks using this read-only kubectl query
inspecting container.googleapis.com/instance_id:
Step 4: Remediation by analyzer category
Load and follow only the reference file that matches the analyzer category from Step 1:
- Path A (workload resource bottlenecks — HBM, host memory, or host CPU utilization): Read Path A: Workload resource bottlenecks [blocked].
- Path B (infrastructure, thermal, ICI, memory bandwidth, or network fabric throttling): Read Path B: Infrastructure or network fabric throttling [blocked].
- Path C (
PERFORMANCE_DEGRADATIONevent fired with all analyzersNOT_DETECTED): Read Path C: Event fired with no DETECTED analyzer [blocked]. - Path D (no
PERFORMANCE_DEGRADATIONevents and steady telemetry): Read Path D: Healthy telemetry [blocked].
Guardrails
- Never change GKE-managed instance groups or VMs through Compute Engine:
Don't run
gcloud compute instance-groups managedcommands, such asdelete, on a node pool's managed instance group. Handle nodes through GKE, as described in Path B: Infrastructure or network fabric throttling [blocked]. - Never cordon nodes for workload resource saturation: If only the HBM capacity, host memory, or host CPU utilization analyzers detected an issue, don't cordon or replace nodes.


