Gke Ai Troubleshooting Tpu Dynamic Slices Monitoring

by google55b4e13eba6dNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Monitors, troubleshoots, and manages GKE TPU Dynamic Slices custom resources. Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, validating single-slice or multi-slice (JobSet) workload manifests, or safely patching stuck finalizers and disabling the slice controller. Don't use for generic GKE cluster node pool creation or standard non-TPU workload management (use gke-basics or gke-cluster-creation instead).

FeaturedIncludes scriptsDevOps & Cloud
AI-generated overview

Monitors and troubleshoots GKE TPU Dynamic Slices, validates workload manifests, and performs risky cleanup operations.

What it does
Guides an agent through inspecting GKE TPU Slice custom resources with kubectl describe, interpreting Status.Conditions lifecycle states, and diagnosing provisioning failures such as node allocation, topology, and reservation block mismatches. It also checks single-slice and multi-slice JobSet workload manifests for required annotations and node selectors. Finally, it provides high-risk resolution procedures for removing stuck finalizers and disabling the slice controller, requiring user confirmation before execution.
When to use it
Use when checking TPU slice lifecycle states, troubleshooting slice provisioning failures, or validating single-slice and multi-slice JobSet workload manifests. It is also intended for safely patching stuck finalizers or disabling the slice controller. It is not for generic GKE node pool creation or standard non-TPU workload management.
Requirements
Requires kubectl and gcloud CLIs configured for the GKE cluster, Cloud Logging enabled for the project, and cluster access. Ships an executable validation script and a reference document.

GKE TPU Dynamic Slices Monitoring & Management

Monitors the status of TPU Slice custom resources, troubleshoots provisioning failures, validates workload manifests on dynamic slices, and performs cleanups.

Prerequisites

  • Cloud Logging enabled for the project.
  • kubectl and gcloud CLIs configured to access the GKE cluster.

Diagnostic Workflow

Step 0: Context Acquisition & Time Window Definition

Gather project, cluster, and slice context using cluster tools or the following parameters:

  • Project ID: {project_id} (e.g., my-gcp-project)
  • Cluster Name: {cluster_name} (e.g., tpu-cluster)
  • Region/Zone: {location} (e.g., us-central1-a)
  • Slice Name: {slice_name} (e.g., test-slice)
  • Issue Time: {timestamp} (Optional; default to the last 30 minutes window [T - 30m] to [T + 30m])

Step 1: Describe the Slice Custom Resource [Low Risk]

When asked to inspect, troubleshoot, or check a slice status, immediately execute kubectl describe slice {slice_name} using available cluster tools to perform the inspection. Parse the resulting Status.Conditions output against the condition table below to diagnose the exact state and provide concrete recommendations.

  • Command:

    bash
    kubectl describe slice {slice_name}
State & Reason Analysis

Analyze the Status.Conditions (especially Type: Ready and its Reason and Status):

Lifecycle State / ReasonMeaningRecommended Action
SliceNotCreatedGKE Slice Controller is initializing the slice and performing resource checks.Wait a few minutes and re-check slice status.
SliceCreationFailedPrerequisites validation failed (e.g., selected nodes don't exist, nodes are already used by another slice, or the topology doesn't match the number of partitions).Verify selected nodes exist, are unallocated, and topology matches partition count.
ACTIVATINGGKE is actively forming and provisioning the TPU slice.Monitor node provisioning.
ACTIVEThe TPU slice is successfully formed and ready to host workloads.Proceed to deploy or check workloads.
ACTIVE_DEGRADEDThe slice is usable, but one or more sub-blocks are degraded.Monitor workload logs for interconnect or device errors. Check faulty node VMs.
FAILEDGKE failed to form the TPU slice (e.g., selected nodes are not part of the same reservation block).Ensure all selected nodes belong to the same reservation block.
DEACTIVATINGThe slice is dismantling (triggered by user deletion or a critical systemic failure).Wait for dismantling to finish, or patch finalizers if stuck.
INCOMPLETEThe terminal phase before the Slice CR is deleted from the cluster.No action required; the resource will be removed shortly.
Provisioning Failure Troubleshooting Checklist

When investigating slice creation or provisioning failures (SliceCreationFailed or FAILED), perform the following verification steps:

  1. Node Existence & Allocation Check: Verify that the selected TPU nodes exist in the cluster and are not already allocated to another slice (kubectl get nodes -l cloud.google.com/gke-tpu-slice, kubectl get slice -A).
  2. Topology Alignment: Confirm that the partition count matches the requested topology dimensions (e.g. topology 2x2 requires 4 nodes).
  3. Reservation Block Alignment Check: Confirm that all selected TPU nodes belong to the same reservation and reservation block.

Step 2: Verify Workload Specification [Low Risk]

Ensure workload manifests are configured correctly to target the dynamic slice.

1. Single-Slice Workload Requirements

Check that the Pod template contains the following annotations and selectors:

  • Annotations:
    • cloud.google.com/gke-tpu-slice-topology: "{topology}" (e.g., "4x4x4")
  • NodeSelector:
    • cloud.google.com/gke-tpu-topology: "{topology}" (e.g., "4x4x4")
    • cloud.google.com/gke-tpu-accelerator: "{accelerator_type}" (e.g., "tpu7x")
    • cloud.google.com/gke-tpu-slice: "{slice_name}" (e.g., "test-slice")
2. Multi-Slice (JobSet) Workload Requirements

If deploying a multi-slice JobSet, verify:

  • JobSet Annotation:
    • alpha.jobset.sigs.k8s.io/exclusive-topology: cloud.google.com/gke-tpu-slice
  • Pod Template Annotations:
    • cloud.google.com/gke-tpu-slice-topology: "{topology}"
  • Pod Template NodeSelector:
    • cloud.google.com/gke-tpu-topology: "{topology}"
    • cloud.google.com/gke-tpu-accelerator: "{accelerator_type}"
    • Note: Do NOT manually specify cloud.google.com/gke-tpu-slice in the nodeSelector; JobSet handles slice assignment automatically.

Resolution & Management Workflow

Resolution 1: Force Delete a Stuck Slice [High Risk]

If a slice is stuck in DEACTIVATING or deletion hangs indefinitely due to stuck finalizers:

  1. Identify Cause: Explain that finalizers on the slice resource (metadata.finalizers) are preventing Kubernetes from completing resource deletion.

  2. Propose Resolution: Propose removing finalizers from the metadata path (/metadata/finalizers) using a JSON patch operation:

    bash
    kubectl patch slice {slice_name} --type json -p='[{"op": "remove", "path": "/metadata/finalizers"}]'
  3. Provide Warning: Explicitly warn the user that removing finalizers bypasses standard controller dismantling and may leave underlying VM, network, or accelerator resources uncleaned or orphaned.

  4. CRITICAL SAFETY MANDATE: The response MUST explicitly ask the user for confirmation (e.g. "Removing finalizers on /metadata/finalizers via JSON patch is a high-risk operation that may leave orphaned resources. Do you confirm you want to apply this patch to slice {slice_name}?") and pause for user confirmation before applying or executing the patch.


Resolution 2: Disable and Clean Up Slice Controller [High Risk]

If dynamic slicing needs to be disabled:

  1. Check for existing Slices:

    bash
    kubectl get slice -A

    Ensure all slices are deleted before disabling the controller.

  2. Disable Slice Controller via gcloud:

    bash
    gcloud container clusters update {cluster_name} \    --location={location} \    --no-enable-slice-controller
  3. Delete the Slice CRD:

    bash
    kubectl delete crd slices.accelerator.gke.io
  4. Clean up Node Labels: Remove GKE TPU Slice labels from all nodes in the cluster:

    bash
    kubectl label nodes --all cloud.google.com/gke-tpu-slice- cloud.google.com/gke-tpu-slice-topology-
  • Safety Rule: Propose the exact commands and confirm before executing disabling or destructive cleanup steps.

Source and attribution

Source:google/skillsinskills/cloud/gke-ai-troubleshooting-tpu-dynamic-slices-monitoringat commit55b4e13

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from google/skills

Dpop Adoption

google

Featured

Implement and debug OAuth 2.0 DPoP (RFC 9449) refresh token sender-constraining for WebCrypto, Node.js ES6, and browser runtimes integrating with Google's OAuth platform. Use when configuring non-extractable asymmetric key pairs (P-256), generating DPoP Proof JWTs for authorization code exchange and token refresh, or handling 400 use_dpop_nonce challenge retry loops at oauth2.googleapis.com/token. Don't use for unconstrained OAuth 2.0 flows (where refresh tokens are not bound to a client key pair), or for Google Cloud IAM / service account authentication.

Awaiting classificationOct 8, 2026

Finding Google Skills

google

Featured

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

Awaiting classificationOct 8, 2026

Spanner Basics

google

Featured

Guides Google Cloud Spanner administration, schema design, querying and performance diagnosis.

Data & AnalyticsOct 8, 2026

Secops Triage

google

Featured

Guides SOC analysts through triaging Google SecOps security alerts, from investigation to closure or escalation.

SecurityOct 8, 2026

Secops Investigate

google

Featured

Guides SOC analysts through deep security incident and entity investigations in Google SecOps using UDM queries and timelines.

SecurityOct 8, 2026

Secops Hunt

google

Featured

Guides proactive threat hunting in Google SecOps using UDM queries, IoC lookback, prevalence and outlier analysis.

SecurityOct 8, 2026