Gke Ai Troubleshooting Tpu Vbar Oom

作者 google55b4e13eba6d無授權條款21K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Diagnoses and prevents vbar_control_agent segfaults, out-of-memory (OOM) errors, and TPU device initialization failures on TPU v6e nodes in GKE caused by race conditions during TPU device resets or high-frequency metrics polling. Use when troubleshooting vbar_control_agent crashes, memory cgroup OOMs in serial console logs, tpu-device-plugin metrics checksum corruption errors, or custom TPU metrics collection conflicts on GKE TPU v6e nodes. Don't use for general non-TPU container OOM troubleshooting or standard GKE node lifecycle operations.

精選包含腳本DevOps & Cloud
AI 產生的概覽

透過 Cloud Logging 查詢診斷並預防 GKE TPU v6e 節點上的 vbar_control_agent 當機與 OOM。

功能
提供逐步診斷流程,用來排查 GKE TPU v6e 節點上 vbar_control_agent 的段錯誤、記憶體 cgroup OOM 以及 TPU 裝置初始化失敗。它提供用於序列埠主控台 OOM 與 tpu-device-plugin 指標總和檢查碼錯誤的 Cloud Logging 篩選範本,並檢查是否有自訂 TPU 指標收集。接著提出解決建議,例如停用自訂指標收集或等待 GKE 韌性更新,並附上檢查清單與故障特徵參考。
適用情境
適用於排查 GKE TPU v6e 節點上的 vbar_control_agent 當機、序列埠主控台記錄中的記憶體 cgroup OOM、tpu-device-plugin 指標總和檢查碼損毀,或自訂 TPU 指標收集衝突。不適用於一般非 TPU 容器 OOM 問題或標準 GKE 節點生命週期操作。
執行需求
專案需啟用 Cloud Logging,並可透過 gcloud 或等效工具存取專案與叢集。隨附驗證指令碼(scripts/validate_queries.sh)與參考檔案;即時診斷使用 query_logs 工具。

TPU Connection Failure and VBAR OOM Troubleshooting

Use this skill to systematically diagnose and prevent vbar_control_agent segfaults and Out-Of-Memory (OOM) errors on TPU v6e nodes.

⚠️ Prerequisites

  • Cloud Logging must be enabled for the project.
  • Access to the project and cluster via gcloud or equivalent tool.

🔍 Diagnostic Workflow

Step 0: Context Acquisition & Time Window Definition

Independently gather required context using available GCP/GKE tools or use the provided {variable} placeholders:

  • {project_id}: The GCP Project ID (e.g., customer-ai-project-123).
  • {cluster_name}: The GKE Cluster Name (e.g., tpu-cluster-prod).
  • {node_name}: The Node Name or Instance ID (e.g., tpu-node-1).
  • {workload_name}: The Workload Name / JobSet Name (e.g., my-training-job-456).
  • {namespace}: The Workload Namespace.
  • {issue_time}: The timestamp of the issue (e.g., 2026-04-14T20:00:00Z).
Time Handling & Execution Rules
  1. Window Calculation: If an issue timestamp {issue_time} is provided, calculate the query time window as [{issue_time} - 30m] to [{issue_time} + 30m].
    • Let {start_time} = {issue_time} - 30m
    • Let {end_time} = {issue_time} + 30m
  2. Informational vs. Live Execution: If the user request is informational or query-formulation (e.g. "How can I check...", "How do I determine..."), or if live GCP project resources are not actively targetable, directly output the calculated time window, log names, and Cloud Logging filter templates without attempting live log execution commands.

Step 1: Check for vbar_control_agent OOMs

Look for specific out of memory messages from vbar_control_agent in serial console logs (serialconsole.googleapis.com%2fserial_port_1_output).

  • Tool to use: query_logs (for live diagnostics)
  • Filter Templates:

Serial Console Logs (OOMs):

sql
logName="projects/{project_id}/logs/serialconsole.googleapis.com%2fserial_port_1_output"AND labels."compute.googleapis.com/resource_name"="{node_name}"AND SEARCH(text_payload, "Memory cgroup out of memory: Killed process .* (vbar_control_ag)")AND timestamp >= "{start_time}"AND timestamp <= "{end_time}"
  • Logic: Presence of Memory cgroup out of memory messages related to vbar_control_agent. Stack traces pointing to libtpu::tpunetd::VBARControlHelper::MetricsReadFromVBAR are a strong indicator.
  • Automation: Proceed to next step automatically after reporting findings.
  • Reference: See references/failure_signatures.md for example log patterns.

Step 2: Investigate tpu-device-plugin Metrics Fetch Failures [Low Risk]

Check if tpu-device-plugin is reporting metric fetch failures.

  • Tool to use: query_logs
  • Filter Template:
sql
resource.type="k8s_container"AND resource.labels.project_id="{project_id}"AND resource.labels.cluster_name="{cluster_name}"AND resource.labels.container_name="tpu-device-plugin"AND severity=ERRORAND textPayload:"metrics fetch failed for .* deviceID and .* device path with error: checksum didn't match with the metrics data. Corrupt data found"AND timestamp >= "{start_time}"AND timestamp <= "{end_time}"
  • Logic: Errors indicating "metrics fetch failed" with "checksum didn't match" suggest vBAR memory corruption.
  • Automation: Proceed to next step automatically after reporting findings.

Step 3: Check for Custom Metrics Collection Usage [Low Risk]

Inspect cluster configurations, workloads, or container specs to determine if custom TPU metrics collection mechanisms are deployed.

  • Action: Check if custom scripts or agents (e.g., using libtpu.sdk.tpumonitoring) are deployed that frequently query GetHostMetrics from vBAR Control Agent.

  • Verification Commands:

    • Kubectl Search (Inspect workload env/specs):
    bash
    kubectl get pods -A -o jsonpath='{range .items[*]}{.metadata.namespace}{"/"}{.metadata.name}{"\t"}{.spec.containers[*].image}{"\n"}{end}'
    • Log Search Filter (query_logs):
    sql
    resource.type="k8s_container"AND resource.labels.project_id="{project_id}"AND resource.labels.cluster_name="{cluster_name}"AND textPayload:"libtpu.sdk.tpumonitoring"AND timestamp >= "{start_time}"AND timestamp <= "{end_time}"
  • Logic: Confirmation of custom metrics collection helps confirm the race condition hypothesis.

🛠️ Resolution Workflow

Resolution 1: Temporarily Disable Custom Metrics Collection [High Risk]

If a custom metrics collection agent is identified, recommend disabling it.

  • Action: Recommend disabling the custom metrics collector.
  • Justification: Prevents reads from vBAR during device resets, stopping crashes and OOMs.

Resolution 2: Await vbar_control_agent Resiliency Update [Low Risk]

Advise that a permanent fix will be available in a future GKE version.

  • Action: Recommend upgrading GKE when the fix is available.
  • Justification: The updated agent will be resilient to memory corruption and gracefully handle reads from unbound vBARs.

📋 copypaste checklist

  • Acquire context and compute [{start_time}, {end_time}] window.
  • Check for vbar_control_agent segfaults and OOMs using query_logs.
  • Investigate tpu-device-plugin failures using query_logs.
  • Inspect for custom metrics collection usage.
  • Advise disabling custom metrics collection if applicable.
  • Advise awaiting resiliency update.

來源與署名

來源:google/skills位於skills/cloud/gke-ai-troubleshooting-tpu-vbar-oom提交55b4e13

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 google/skills 的技能

Dpop Adoption

google

精選

指導為 Google OAuth 平台實作 OAuth 2.0 DPoP(RFC 9449)傳送方約束的更新權杖。

Security21K今天更新

Finding Google Skills

google

精選

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

待分類21K今天更新

Spanner Basics

google

精選

指導 Google Cloud Spanner 的執行個體與資料庫管理、結構定義設計、查詢與效能診斷。

Data & Analytics21K今天更新

Secops Triage

google

精選

引導 SOC 分析師對 Google SecOps 安全警示進行分診,從調查到結案或升級。

Security21K今天更新

Secops Investigate

google

精選

指導 SOC 分析師在 Google SecOps 中使用 UDM 查詢與時間軸進行深入的安全事件與實體調查。

Security21K今天更新

Secops Hunt

google

精選

指導在 Google SecOps 中使用 UDM 查詢、IoC 回溯、普遍性與異常分析進行主動威脅狩獵。

Security21K今天更新