Gke Ai Troubleshooting Tpu Performance Degradation

google/skills/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation

by googlec6c7e67107f02e3423df8481272f296bbe61902aNo license21K starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated today

Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system metrics (`kubernetes.io/node/accelerator/*`). Distinguishes hardware and network fabric throttling from workload resource bottlenecks (HBM capacity, host memory, or host CPU saturation). Use when TPU training throughput or duty cycle drops without crashing pods, when `PERFORMANCE_DEGRADATION` monitored events fire, or when triaging slow multi-slice training steps. Don't use for complete multi-slice XLA execution stalls with `HANG_DETECTED` logs (use gke-ai-troubleshooting-tpu-mxla-hang) or pod eviction/interruption restarts (use gke-ai-troubleshooting-jobset-interruption).

FeaturedInstructions onlyDevOps & Cloud

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/path-a-workload-bottlenecks.md2.3 KBtext/markdown
references/path-b-infrastructure-throttling.md2.2 KBtext/markdown
references/path-c-no-detected-analyzer.md1.5 KBtext/markdown
references/path-d-healthy-telemetry.md323 Btext/markdown
SKILL.md13.6 KBtext/markdown

Source and attribution

Source:google/skillsinskills/cloud/gke-ai-troubleshooting-tpu-performance-degradationat commitc6c7e67

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal