GKE Cluster Creation
This reference guides creating Google Kubernetes Engine (GKE) clusters by providing a set of best-practice templates and guiding through mode selection and customization. The golden path Autopilot configuration is the default for all new clusters.
MCP Tools:
list_clusters,create_cluster,get_cluster,list_operations,get_operation
Workflow
- Discover context: Use
list_clustersto see existing clusters. Usegcloud config get-value projectif project unknown. - Gather inputs:
project_id,location(region or zone),cluster_name, environment type. If missing essential details, ask the user before taking action. - Select mode & explain trade-offs: If the user hasn't specified a template or mode, present the available templates (e.g., Autopilot, Standard Regional, GPU Inference, AI Hypercompute) and explain key trade-offs (Cost vs. Availability, Autopilot vs. Standard node management).
- Configure networking: auto-create subnet (default) or bring-your-own.
- Review golden path settings: present the default configuration block
(
gcloudcommand orcreate_clusterJSON payload) and confirm with the user before creation. - Create: Use MCP
create_clustertool orgcloudCLI. - Track: Use
get_operationto monitor creation progress. - Verify: Use
get_clusterwithreadMask="*"to confirm golden path settings applied.
Mode Selection
Rule: Default to Autopilot unless the customer has a specific requirement that Autopilot cannot satisfy.
Best Practices
When guiding the user or generating configurations, adhere to these GKE best practices:
Security & Networking
- Private Clusters: Default to private clusters (
enablePrivateNodes: true) with a private control plane and restricted public endpoints (enable-master-authorized-networks) to minimize attack surface. - VPC-Native Networking: Use VPC-native clusters (
useIpAliases: true/--enable-ip-alias) to enable alias IP ranges and pod-level firewall rules. - Workload Identity: Prefer Workload Identity (
workloadPool: <PROJECT_ID>.svc.id.goog) for securely granting GKE workloads access to Google Cloud services instead of static service account keys. - Shielded GKE Nodes: Enable Shielded GKE Nodes
(
--enable-shielded-nodes,--enable-secure-boot) against rootkits and bootkits. - Least Privilege (RBAC): Institute strict Role-Based Access Control
limits (
scoped-rbs-bindings).
Cost Optimization
- Autoscaling: Enable Cluster Autoscaler and Horizontal/Vertical Pod
Autoscaler (
--enable-autoscaling,--enable-vertical-pod-autoscaling) to adjust resources based on demand. - Right-Sizing & Spot VMs: Choose appropriate machine types and node
counts. Consider Spot VMs (
--spot) for fault-tolerant, non-critical batch or inference workloads.
High Availability & Reliability
- Regional Clusters: Use Regional Clusters for production environments to
ensure control plane replication across multiple zones (
--regioninstead of--zone). Note: Standard regional creates nodes across 3 zones by default. - Pod Disruption Budgets: Recommend setting Pod Disruption Budgets for application stability during node maintenance.
- Release Channels: Subscribe to a release channel (
REGULARorSTABLE) for automated, safer cluster upgrades.
Templates
1. Golden Path Autopilot (Production)
This is the default. All settings match
../gke-golden-path/assets/golden-path-autopilot.yaml.
Via gcloud:
Via MCP (create_cluster):
2. Autopilot Dev/Test
Relaxes some golden path defaults for cost savings and easier access in non-production.
Via gcloud:
Via MCP (create_cluster):
Warning: This does not apply golden path security hardening. Suitable for dev/test only.
3. Standard Regional (High Availability / Custom Requirements)
Best when Autopilot cannot be used (e.g., custom kernel tuning, specific node OS requirements). Creates 3 nodes across zones by default.
Via gcloud:
Via MCP (create_cluster):
4. GPU Inference & AI Workloads (L4 / ComputeClass)
Best for: AI/ML Inference, small model serving. Can be provisioned via
Autopilot + ComputeClass or via Standard node pool with g2-standard-4
(nvidia-l4). Note: Requires g2-standard-4 quota.
Autopilot ComputeClass / GIQ approach:
Standard Node Pool approach via MCP (create_cluster):
5. AI Hypercompute (A3 HighGPU / Large Model Serving)
Best for: Large-scale LLM / AI model training and hypercompute inference. Note:
High hourly cost and strict quota requirements (a3-highgpu-8g /
nvidia-h100-80gb-hbm3).
Via gcloud:
Via MCP (create_cluster):
Instructions
- ALWAYS ask for
project_idif not in context. - ALWAYS ask for
region(or location). - ALWAYS ask for a unique
cluster_name. - DEFAULT to golden path Autopilot unless customer specifies otherwise or has custom node/kernel/hypercompute requirements.
- ALWAYS WARN when deviating to GKE Standard, highlighting that it deviates from the golden path and explaining the added operational/management overhead (manually managing node pools, upgrades, and autoscaling).
- EXPLAIN TRADE-OFFS when presenting templates or mode choices to the user if they haven't specified one (e.g., Autopilot vs Standard, Cost vs Availability).
- PRESENT THE CONFIGURATION block (
gcloudcommand or JSON payload) and ask for confirmation before calling any creation tool. - WARN about Day-0 decisions (networking, private nodes) that are hard to change later.
- WARN explicitly about cost and quota requirements when the user selects
GPU (
g2-standard-4,a3-highgpu-8g), TPU, or multi-region/regional clusters (--regiondefaults to 3 zones). - When using MCP
create_cluster, thecluster.nameparameter should be the short name (e.g.,my-cluster), not the full resource path (projects/<PROJECT_ID>/locations/<REGION>/clusters/<CLUSTER_NAME>). Theparentparameter defines the scope (projects/<PROJECT_ID>/locations/<REGION>).
