Gke Productionize

by google55b4e13eba6dNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Orchestrates comprehensive production readiness reviews and assessments for GKE clusters and workloads across scalability, security, reliability, observability, backup/DR, and cost optimization. Use when asked to productionize, prepare, assess, audit, or review a GKE cluster or workload before going live to production. Don't use for deep-dive single-domain implementation (use specific domain skills like gke-workload-scaling, gke-platform-security, gke-workload-security, gke-service-networking, gke-reliability instead).

FeaturedInstructions onlyDevOps & CloudSecurity
AI-generated overview

Orchestrates GKE production readiness reviews across scalability, security, reliability, observability, backup and cost.

What it does
This orchestrator skill guides a production readiness assessment of a GKE cluster or workload. It runs discovery commands to inspect cluster and workload configuration, then delegates to specialized domain skills for scaling, observability, reliability, security, backup/DR, networking, cost, upgrades and golden-path defaults. It produces a per-domain RAG status report with a weighted overall readiness score and verdict.
When to use it
Use when asked to productionize, prepare, assess, audit or review a GKE cluster or workload before going live. It is intended for broad readiness reviews rather than deep single-domain implementation, which should use the specific domain skills instead.
Requirements
Requires gcloud and kubectl access to the target GKE cluster and project, plus the referenced specialized GKE skills. Ships no scripts; instructions only.

GKE Productionize Skill

This skill acts as a high-level orchestrator for preparing a GKE cluster and its workloads for production readiness.

[!IMPORTANT] This is a meta-skill or orchestrator skill. You are expected to invoke and run many other specialized skills listed in this document as part of the overall productionization process. Do not attempt to implement all production readiness features directly within this skill; instead, use this skill to assess the environment and then delegate to the specific skills for each domain.

Scope

This skill is adaptable to:

  • A single application (already on Kubernetes or not).
  • A set of applications.
  • A target cluster.

Workflow

1. Discovery Phase

Before making recommendations, discover the current state of the environment.

Cluster Discovery

Run these commands to understand the cluster setup:

  • Check cluster details: gcloud container clusters describe {cluster_name} --location {location} --project {project}

  • Check for Autopilot vs Standard: Look for the following block in the describe output:

    yaml
    autopilot:  enabled: true
  • Check release channel: Look for releaseChannel.

Workload Discovery

If a specific application is targeted, discover its configuration:

  • Get deployment/statefulset details: kubectl get deployment {app_name} -n {namespace} -o yaml
  • Check for dedicated namespace and labels: kubectl get namespace {namespace} -o yaml (Look for Pod Security Standards labels).
  • Check for dedicated service account usage: kubectl get pods -n {namespace} -o custom-columns="NAME:.metadata.name,SERVICE_ACCOUNT:.spec.serviceAccountName"
  • Check for resource requests and limits.
  • Check for liveness, readiness, and startup probes.
  • Check for HPA: kubectl get hpa -n {namespace}
  • Check for PDB: kubectl get pdb -n {namespace}
  • Check for NetworkPolicies: kubectl get networkpolicy -n {namespace}

2. Production Readiness Assessment

Before implementation, you MUST run the skills for each relevant specialized area listed below and incorporate its guidance into your assessment and plan. Failure to do so will result in a non-compliant production configuration.

A. App Onboarding (Pre-Kubernetes)

If the application is not yet running on GKE, you MUST run the gke-app-onboarding skill for planning containerization, image building, and basic deployment.

B. Scalability & Resource Management

Ensure workloads have appropriate resources and autoscaling.

  • Action: You MUST run the gke-workload-scaling skill for configuring HPA, VPA, and resource limits.
C. Observability

Ensure adequate logging and monitoring are in place.

  • Action: You MUST run the gke-observability skill for setting up Cloud Logging, Monitoring, and Managed Prometheus.
D. Reliability

Ensure high availability and graceful degradation.

  • Action: You MUST run the gke-reliability skill for configuring regional clusters, PDBs, and health probes.
E. Security

Harden the cluster and workloads.

  • Action: You MUST run the gke-platform-security and gke-workload-security skills for Workload Identity, Network Policies, and Shielded Nodes.
  • Namespace Isolation: Ensure workloads run in dedicated namespaces with Pod Security Standards (PSS) enforced via labels.
  • Least Privilege: Ensure workloads use dedicated ServiceAccounts instead of the default ServiceAccount.
F. Backup & Disaster Recovery

Ensure stateful data is protected.

  • Action: You MUST run the gke-backup-dr skill for configuring Backup for GKE and restore procedures.
G. Edge Security & Ingress

Secure external access.

  • Action: You MUST run the gke-service-networking skill for Gateway API, Ingress, and Cloud Armor.
H. Cost Optimization

Ensure efficient use of resources.

  • Action: You MUST run the gke-cost-optimization skill for strategies on rightsizing, quotas, and Spot VMs.
I. Upgrades & Maintenance Posture

Ensure a safe, predictable upgrade posture.

  • Action: You MUST run the gke-upgrades skill for release channel selection, maintenance windows/exclusions, and node pool upgrade strategy.
J. Golden Path Defaults Audit

Ensure the cluster configuration matches recommended defaults.

  • Action: You MUST run the gke-golden-path skill to compare the cluster against golden path defaults and report deviations with severity and remediation.

3. Production Readiness Scoring

After the assessment, provide a summary report with a RAG (Red, Amber, Green) status for each area and an overall readiness score. This helps prioritize remediation efforts.

Apply this rubric deterministically so repeated assessments of the same environment produce the same result:

  1. Per-domain criteria: For each assessed domain (A-J), list the concrete checks performed (from the domain skill's guidance) and classify each check as pass, fail-critical (production-blocking, e.g., no resource requests, no backups for stateful data, public control plane in a locked down environment), or fail-minor (improvement, e.g., missing VPA recommendations, no Spot usage for batch).
  2. RAG mapping (per domain):
    • Red = one or more fail-critical checks.
    • Amber = no fail-critical, but one or more fail-minor checks.
    • Green = all checks pass.
  3. Domain score: Green = 100, Amber = 50, Red = 0.
  4. Weighted overall score: weight Security, Reliability, and Backup/DR at 2x; all other assessed domains at 1x. Overall score = sum(domain score x weight) / sum(weights), rounded to the nearest integer. Exclude domains that are not applicable (e.g., Backup/DR for fully stateless workloads) from both sums and note the exclusion.
  5. Readiness verdict: >= 90 with no Red domains = "Production ready"; 70-89 with no Red domains = "Ready with follow-ups"; anything else = "Not production ready".

In the report, show the per-domain check lists, RAG status, weights, and the computed overall score.

Adaptability Guidelines

  • Single App: Focus on Health Probes, HPA, Resource Limits, PDB, and Workload Identity for that specific app.
  • Cluster Wide: Focus on Cluster Autoscaler, Multi-zonal setup, Release Channels, Maintenance Windows, and default Network Policies.
  • Proactive Execution: Proactively execute relevant skills (e.g., observability, security, scaling, reliability) to assess and propose improvements, seeking user confirmation before applying state-changing implementations.

Source and attribution

Source:google/skillsinskills/cloud/gke-productionizeat commit55b4e13

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from google/skills

Dpop Adoption

google

Featured

Implement and debug OAuth 2.0 DPoP (RFC 9449) refresh token sender-constraining for WebCrypto, Node.js ES6, and browser runtimes integrating with Google's OAuth platform. Use when configuring non-extractable asymmetric key pairs (P-256), generating DPoP Proof JWTs for authorization code exchange and token refresh, or handling 400 use_dpop_nonce challenge retry loops at oauth2.googleapis.com/token. Don't use for unconstrained OAuth 2.0 flows (where refresh tokens are not bound to a client key pair), or for Google Cloud IAM / service account authentication.

Awaiting classificationOct 8, 2026

Finding Google Skills

google

Featured

Google platform decision and setup guidance, loaded on demand from Google's skill catalog. Use when a developer is choosing or setting up part of their stack, such as where to run a service, a database, storage, messaging, authentication, analytics, ads, or AI model serving, and a Google product is a reasonable candidate - whether or not a vendor is named - or when a request names a Google product or API. Brings in the matching Google skill so the answer can weigh Google options, their trade-offs, and when they are not the right fit. Skip when the stack is already settled on another provider and no Google product is named, or the task involves no platform choice.

Awaiting classificationOct 8, 2026

Spanner Basics

google

Featured

Assists in provisioning instances and databases, designing performant schemas, and querying data in Spanner. Use when designing primary keys, writing SQL queries or client library code, or diagnosing performance issues.

Awaiting classificationOct 8, 2026

Secops Triage

google

Featured

Guides SOC analysts through triaging Google SecOps security alerts, from investigation to closure or escalation.

SecurityOct 8, 2026

Secops Investigate

google

Featured

Guides SOC analysts through deep security incident and entity investigations in Google SecOps using UDM queries and timelines.

SecurityOct 8, 2026

Secops Hunt

google

Featured

Expert guidance for proactive threat hunting in Google SecOps. Use when proactively hunting for threats, retroactively analyzing indicators of compromise (IoCs), performing prevalence searches across enterprise events, hunting for MITRE ATT&CK techniques, or detecting behavioral and statistical outliers using UDM queries. Don't use for incoming alert triage (use secops-triage), active incident response and timeline deep-dives on a known breach (use secops-investigate), or detection rule authoring (use secops-detection-engineering).

Awaiting classificationOct 8, 2026