Gke Ai Troubleshooting Node Unresponsive Timeout

by google55b4e13eba6dNo license21K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Diagnose and mitigate GKE TPU or GPU nodes stuck in NotReady / NodeStatusUnknown ("Kubelet stopped posting node status") due to host kernel panics, hardware lockups, or disabled node auto-repair. Use when nodes stop heartbeating beyond the node auto-repair threshold and pods remain stuck in Terminating. Don't use for healthy nodes, pod-only application crashes, or routine GKE upgrades.

FeaturedInstructions only

Troubleshoot unresponsive GKE TPU and GPU nodes (NodeStatusUnknown)

When the Compute Engine host of a TPU or GPU node has a fatal hardware error, kernel panic, or non-maskable interrupt (NMI) lockup, the guest OS stops responding. The kubelet can no longer send heartbeats, so the node Ready condition becomes Unknown with Reason: NodeStatusUnknown (Kubelet stopped posting node status.). If node auto-repair is disabled on the node pool, GKE doesn't repair the node. The node can stay NotReady, and pods on it can stay in Terminating, which blocks multi-host JobSet workloads from recovering.

Prerequisites

  • Tools: Install the Google Cloud SDK (gcloud) and kubectl.
  • Cloud Billing & Project Configuration: Verify an active billing account is linked (gcloud billing projects describe {project_id}), authenticate (gcloud auth login), set the target project (gcloud config set project {project_id}), and ensure container.googleapis.com, compute.googleapis.com, logging.googleapis.com, and monitoring.googleapis.com are enabled.
  • Required IAM Roles:
    • Kubernetes Engine Viewer (roles/container.viewer)
    • Compute Viewer (roles/compute.viewer)
    • Logs Viewer (roles/logging.viewer)
    • Monitoring Viewer (roles/monitoring.viewer)
    • For remediation ([High Risk] steps): Kubernetes Engine Cluster Admin (roles/container.clusterAdmin)
  • Documentation:

Read-only rule: Run read-only diagnostic commands only. Never drain, delete, or re-create nodes, or run any other command that changes the cluster. Give the user any fix to apply themselves.

When you recommend a fix, link the doc section that describes it.


Diagnostic workflow

Step 0: Collect context and set the investigation window [Low Risk]

Collect the target parameters. By default, query a 60-minute window `[T - 30m, T

  • 30m]around{issue_time}`:
  • {project_id}: Google Cloud project ID
  • {cluster_name}: GKE cluster name
  • {location}: Cluster region or zone
  • {nodepool_name}: Target TPU or GPU node pool name
  • {node_name}: Unresponsive GKE node name (and its Compute Engine {zone})
  • {issue_time}: Incident timestamp in RFC3339 UTC
  • {start_time}: {issue_time} - 30m
  • {end_time}: {issue_time} + 30m

Step 1: Verify the NodeStatusUnknown heartbeat timeout [Low Risk]

  1. Check Kubernetes node conditions: To inspect the node status and verify whether the Ready condition is Unknown with Reason: NodeStatusUnknown (Kubelet stopped posting node status.), follow the instructions in the section Check the node's status and conditions.
  2. Query Cloud Logging (read-only LQL): Query k8s_node and k8s_cluster logs across [{start_time}, {end_time}] to confirm when the control plane lost heartbeat contact with {node_name}:
(resource.type="k8s_node" OR resource.type="k8s_cluster")resource.labels.cluster_name="{cluster_name}"("{node_name}" AND ("NodeNotReady" OR "NodeStatusUnknown" OR "Kubelet stopped posting node status"))timestamp >= "{start_time}" AND timestamp <= "{end_time}"
  1. Query Cloud Monitoring (read-only PromQL): Correlate the duration of the Unknown state using the GKE system metric kubernetes.io/node/status_condition (kubernetes_io:node_status_condition, GKE 1.32.1-gke.1357001+) documented in Monitor health metrics for TPU nodes and node pools, filtered by condition="Ready" and status="Unknown":
promql
kubernetes_io:node_status_condition{  monitored_resource="k8s_node",  cluster_name="{cluster_name}",  node_name="{node_name}",  condition="Ready",  status="Unknown"}
  • Decision logic:
    • If the node Ready condition is True and NodeStatusUnknown is absent, rule out an unresponsive node timeout and pivot to workload-level troubleshooting (for example, Troubleshoot OOM events) rather than repairing or draining the node.
    • If Ready is Unknown (NodeStatusUnknown), proceed to Step 2.

Step 2: Inspect serial port output for a kernel panic [Low Risk]

The guest OS can't send logs after a fatal kernel freeze, so check the serial port output of the node's VM:

  • Cloud Logging: If serial port logging is enabled (i.e. VM metadata serial-port-logging-enable is set to true), GKE system logs in Cloud Logging include the node's serial port output. See the section System logs. Use Cloud Logging when the VM is stopped or has already been replaced by auto-repair, or when you need more than the most recent output.
  • Running VM: Follow Viewing serial port output to retrieve the serial port 1 output (gcloud compute instances get-serial-port-output with --port=1) for {node_name} in {zone}. This method returns only the most recent 1 MB of output per port.

Consult Troubleshoot Linux VM boot issues due to kernel panic to identify documented kernel panic and hardware crash patterns (such as Fatal Machine check, hung_task: blocked tasks, or NMI: Not continuing) in the serial port output.


Step 3: Check node auto-repair status [Low Risk]

Find out why GKE hasn't repaired the unresponsive node:


Step 4: Check Compute Engine system events [Low Risk]

To list the system events for {node_name} around {issue_time}, follow the instructions in the section Querying Cloud Audit Logs. Compare the method field with the table in the "Reviewing Cloud Audit Logs" section of the same document, and look for:

  • compute.instances.hostError: a hardware or software issue on the physical host caused the VM to crash.
  • compute.instances.preempted: Compute Engine preempted a Spot VM or preemptible VM. For preempted nodes, also see Confirm node preemption.
  • compute.instances.automaticRestart: Compute Engine restarted the VM after a hostError or terminateOnHostMaintenance event.
  • compute.instances.guestTerminate: the VM's operating system initiated the shutdown.

Step 5: Resolution [High Risk]

Guardrails:

  • Never force-delete stuck Terminating pods on an unresponsive node. Force deletion doesn't wait for the kubelet to confirm that the pod has stopped, so a replacement pod can start while the old one is still running. See Force Delete StatefulSet Pods.
  • Never delete GKE-managed Compute Engine VM instances directly (gcloud compute instances delete). Instead, check how long the node has reported NodeStatusUnknown (Step 1), check the serial console output (Step 2), and rely on node auto-repair, as described in the following steps.
  1. Enable node auto-repair (Standard clusters only):
  2. Let GKE repair the node:
    • Explain that GKE repairs a node that reports NotReady or no status for the documented time threshold by draining and re-creating it, and link the "Repair criteria" and "Node repair process" sections of Auto-repair nodes. GKE waits one hour for the drain to complete. If the drain doesn't complete, GKE shuts the node down and creates a new node. Tell the user to expect this one-hour drain wait before GKE re-creates the node.
    • For multi-host TPU slice node pools, note per Node auto repair in TPU slice nodes that the entire node pool is re-created.
  3. Verify node recovery:
    • After the repair completes, follow the instructions in the section Verify that the node has recovered to verify that the repaired node returns to Ready status, then confirm that the JobSet pods are running again.

Source and attribution

Source:google/skillsinskills/cloud/gke-ai-troubleshooting-node-unresponsive-timeoutat commit55b4e13

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from google/skills

Dpop Adoption

google

Featured

Guides implementation of OAuth 2.0 DPoP (RFC 9449) sender-constrained refresh tokens for Google's OAuth platform.

Security21Kupdated today

Finding Google Skills

google

Featured

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

Awaiting classification21Kupdated today

Spanner Basics

google

Featured

Guides Google Cloud Spanner administration, schema design, querying and performance diagnosis.

Data & Analytics21Kupdated today

Secops Triage

google

Featured

Guides SOC analysts through triaging Google SecOps security alerts, from investigation to closure or escalation.

Security21Kupdated today

Secops Investigate

google

Featured

Guides SOC analysts through deep security incident and entity investigations in Google SecOps using UDM queries and timelines.

Security21Kupdated today

Secops Hunt

google

Featured

Guides proactive threat hunting in Google SecOps using UDM queries, IoC lookback, prevalence and outlier analysis.

Security21Kupdated today
Gke Ai Troubleshooting Node Unresponsive Timeout Agent Skill | SourceWeft