GKE AI/ML Inference
Routing Note: For migrating existing AI workloads to GKE, open
google-cloud-solution-guided-gke-ai-migration/SKILL.md. For GKE RAG with Cloud SQL or AlloyDB (pgvector), opengoogle-cloud-solution-rag-enterprise-search-gke-sqldb/SKILL.md.
This reference covers deploying AI/ML inference workloads on GKE using Google's Inference Quickstart (GIQ) and best practices for LLM serving.
MCP Tools:
apply_k8s_manifest,get_k8s_resource,get_k8s_logs,get_k8s_rollout_status,describe_k8s_resource,list_k8s_events. CLI-only:gcloud container ai profiles *
When to Use
- Deploy an AI model (Llama, Gemma, Mistral, etc.) to GKE
- Generate optimized Kubernetes manifests for inference
- Select GPU/TPU accelerators for model serving
- Configure autoscaling for LLM inference
Prerequisites
- A golden path GKE Autopilot cluster (GPU workloads are supported via ComputeClasses and NAP)
gcloudCLI authenticated- Sufficient GPU/TPU quota in the target region
Workflow
1. Discovery: Find Models and Hardware
2. Generate Manifest
Parameters:
--model: Model ID (e.g.,gemma-2-9b-it,llama-3-8b)--model-server: Inference server (vllm,tgi,triton,tensorrt-llm)--accelerator-type: GPU/TPU type (nvidia-l4,nvidia-tesla-a100,nvidia-h100-80gb)--target-ntpot-milliseconds: Target Normalized Time Per Output Token (optional, for latency optimization)
Example:
3. Review and Deploy
Some models require Hugging Face tokens. Create a Kubernetes Secret and reference it in the manifest.
GPU ComputeClass for Inference
For Autopilot clusters, create a ComputeClass to target GPU nodes:
Accelerator Selection Guide
Autoscaling LLM Inference
GPU-based autoscaling
Use custom metrics for GPU utilization:
Best practices for inference autoscaling
- Use DCGM metrics: Golden path enables DCGM monitoring for GPU utilization metrics
- Set appropriate minReplicas: At least 1 for always-on serving; 0 for batch/on-demand
- Tune scale-down delay: LLM model loading is slow; use longer stabilization windows
- Consider queue depth: Scale on pending requests rather than pure GPU utilization for latency-sensitive workloads
Optimization Tips
- Quantization: Use quantized models (GPTQ, AWQ) to reduce GPU memory and increase throughput
- Batching: Configure model server batch size for throughput vs latency trade-off
- Tensor parallelism: Split large models across multiple GPUs within a node
- KV cache optimization: Tune
--gpu-memory-utilizationin vLLM for KV cache allocation
