
Gke Ai Troubleshooting Tpu Performance Degradation
google/skills/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation作者 googlec6c7e67107f02e3423df8481272f296bbe61902a无许可证21K 个星标收录于 2026年10月9日更新于 2026年10月9日仓库今天更新
Diagnose GKE Cloud TPU training throughput drops and step-time regressions (15%+ TPU duty-cycle drop) using ML Diagnostics Workload Monitoring (`gcloud alpha mldiagnostics monitored-events` / `hypercomputecluster.googleapis.com/v1alpha`) and 1-minute Cloud Monitoring system metrics (`kubernetes.io/node/accelerator/*`). Distinguishes hardware and network fabric throttling from workload resource bottlenecks (HBM capacity, host memory, or host CPU saturation). Use when TPU training throughput or duty cycle drops without crashing pods, when `PERFORMANCE_DEGRADATION` monitored events fire, or when triaging slow multi-slice training steps. Don't use for complete multi-slice XLA execution stalls with `HANG_DETECTED` logs (use gke-ai-troubleshooting-tpu-mxla-hang) or pod eviction/interruption restarts (use gke-ai-troubleshooting-jobset-interruption).
添加到 SourceWeft 工作区
- 在控制台中打开该技能,并将它添加到工作区。
- 为需要使用它的对话启用该技能。
该技能仅含说明:不附带任何可执行的脚本。
添加到 SourceWeft系统会先要求你登录,然后直接带你回到这个技能。
让你的智能体来安装
把这段提示词粘贴到 Claude Code、Codex、Cursor 或其他能运行命令的智能体中,也可以粘贴到 SourceWeft 对话里。智能体会阅读这个技能的安装说明,向你展示它的来源、许可证和脚本情况,在你同意后用 SourceWeft CLI 安装。
阅读 https://sourceweft.com/skills/gh-google-skills-gke-ai-troubleshooting-tpu-performance-degradation-skills-cloud-gke-ai-troubleshooting-tpu-performance-degradation-f347f1c211b2fa74/install.md 中的说明,按说明安装这个技能。安装前先告诉我它的来源、许可证以及是否附带脚本,等我确认。修改我电脑上的其他任何内容之前也要先问我。用命令行自行安装
适用于 Claude Code、Codex、Cursor 及其他本地智能体。SourceWeft CLI 会从源代码仓库中获取此处扫描过的那次提交,并根据扫描时记录的哈希值逐一校验每个文件。只要有任何不一致,就不会写入任何内容。
npx @sourceweft/cli skills install @google/gke-ai-troubleshooting-tpu-performance-degradation添加 --agent claude-code、codex、cursor 或 universal 来选择安装给哪个智能体(默认为 Claude Code)。
上游安装器——未经 SourceWeft 校验
开源的 skills 安装器会获取同一个固定的提交,但不会根据 SourceWeft 记录的哈希值校验文件。
npx skills add https://github.com/google/skills/tree/c6c7e67107f02e3423df8481272f296bbe61902a/skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation来源与署名
来源:google/skills位于skills/cloud/gke-ai-troubleshooting-tpu-performance-degradation提交c6c7e67
许可证: 无许可证
内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。
更多来自 google/skills 的技能

Sign In With Google Web
为各类 Web 架构提供 Sign In With Google(GIS)集成与安全实现的指导。

Gke Ai Troubleshooting Tpu Mxla Hang
使用 ML Diagnostics MXLA 挂起分析器报告和指标诊断 GKE Cloud TPU 多切片训练挂起。

Google Cloud Scc Remediation
指导修复 Google Cloud Security Command Center 发现项,提供需用户批准的方案与参考手册。

Google Cloud Recipe Onboarding
引导开发者完成首次 Google Cloud 上手:身份验证、项目设置、结算关联与验证。

Google Cloud Recipe Auth
指导用户、服务账号和工作负载如何对 Google Cloud 服务与 API 进行身份验证和授权。

Gcloud
为在 Google Cloud 上执行 gcloud CLI 命令提供安全护栏、语法校验和数据精简规则。
更多DevOps & Cloud技能

Update Screenshots
microsoft
在 Screenshots & Tests 检查失败后,更新已提交的 CI 截图哈希基线。

Playwright Devops
microsoft
面向 Playwright 的 DevOps 工作流:分析 main 分支最新提交的 GitHub Actions 失败并下载失败作业日志。

M5 Onboard
anthropics
通过 USB 检测 M5Stack ESP32 开发板,刷写 UIFlow 2.0 固件并安装 MicroPython 应用包。

Runbook
anthropics
为重复性任务创建或更新分步运维手册,包含故障排查、回滚和升级流程。

Incident Response
anthropics
指导事故响应流程:严重级别判定、状态更新、缓解跟踪以及无责事后复盘。

Deploy Checklist
anthropics
生成部署前就绪检查清单,涵盖部署前、部署、部署后和回滚触发条件。