Gcp Pipeline Resource Provisioning

by gemini-cli-extensions2df10e25bbf7Apache-2.0215 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Automates declarative resource creation and provisioning for data pipelines, supporting BigQuery, Dataform, Dataproc, BigQuery Data Transfer Service (DTS), and other resources. It manages environment-specific configurations (dev, staging, prod) through a deployment.yaml file. Use when: - Modifying or creating deployment.yaml for deployment settings. - Resolving environment-specific variables (e.g., Project IDs, Regions) for deployment. - Provisioning supported infrastructure like BigQuery datasets/tables, Dataform resources, or DTS resources via deployment.yaml. Do not use when: - Resources already exist. - Managing resources not supported by `gcloud beta orchestration-pipelines resource-types list`. - Managing general cloud infrastructure (VMs, networks, Kubernetes, IAM policies), which are better suited for Terraform. - Infrastructure spans multiple cloud providers (AWS, Azure, etc.). - Already uses Terraform for the target resources.

Instructions onlyDevOps & Cloud
AI-generated overview

Generates and deploys a deployment.yaml that provisions GCP data pipeline resources across environments.

What it does
This skill guides an agent through creating or updating a deployment.yaml file in the repository root that declares GCP data pipeline resources such as BigQuery datasets and tables, Dataform resources, Dataproc, and BigQuery Data Transfer Service transfers. It covers discovering environment parameters with gcloud and git commands, resolving environment-specific values for dev, stage, and prod, referencing secrets from Secret Manager, validating the configuration, and running the deployment command. It produces the deployment.yaml file and an executed deployment.
When to use it
Use it when provisioning or modifying supported GCP pipeline resources through deployment.yaml, or when resolving environment-specific variables such as project IDs and regions. It is not intended for resources that already exist, general cloud infrastructure like VMs, networks, Kubernetes, or IAM, multi-cloud setups, or targets already managed with Terraform.
Requirements
Requires the gcloud CLI with the beta orchestration-pipelines commands, an authenticated GCP project, git for the commit SHA, and access to Secret Manager for secret references. It ships no scripts; it relies on a reference spec file and user-provided environment names.

How to use this skill

Create or update existing deployment.yaml file and deploy resources. All configuration files MUST be maintained together in the repository root.

Mandatory labels

[!IMPORTANT]

Whenever you generate resource definitions in deployment.yaml, you MUST directly populate the datacloud label under definition.labels for every resource to track the source of creation. Determine the value based on your current IDE environment:

  • For Antigravity, set datacloud: "antigravity"
  • For VS Code, set datacloud: "vscode"
  • For any other environment, set datacloud: "other"

Do not use a variable substitution for this label; hardcode the appropriate string value directly into each resource definition (e.g., replacing __REQUIRED_LABEL__ placeholders).

Special rule for BigQuery DTS Ingestion: Whenever you generate a bigquerydatatransfer.transferConfig in deployment.yaml, you MUST also explicitly define its target destination bigquery.dataset in the same file and apply the datacloud label to it. You must do this even if the dataset already exists, to ensure the destination dataset's labels are patched and updated.

Step 1: Supported Resource Types

The framework supports deploying various GCP resources. To see the comprehensive list of supported resource types, run the following command:

gcloud beta orchestration-pipelines resource-types list

Refer to: references/gcp-pipeline-resource-provisioning_spec.md to understand the template for deployment.yaml.

Step 2: Discover Environment Parameters

Before generating configurations, discover the actual values for the target project, region, environment, and commit SHA.

[!TIP]

If deployment.yaml already exists in the repository root, prioritize extracting project and region from the target environment configuration (e.g., dev).

  1. Project ID:

    bash
    gcloud config get project
  2. Project Number:

    bash
    gcloud projects describe $(gcloud config get project) --format="value(projectNumber)"
  3. Region:

    bash
    gcloud config get-value compute/region
  4. Commit SHA:

    bash
    git rev-parse HEAD
  5. Environment Name: If initialization is needed, you MUST ask the user for the environment name. If the user does not provide it, use dev as the default.

[!TIP]

Use these commands to replace placeholders like YOUR_PROJECT_ID with actual values. Always remove associated comments that start with TODO once replaced.

Step 3: Generate or update deployment.yaml

Create or update deployment.yaml in the repository root. This file maps supported environments (dev, stage, prod) to their specific configurations and resources.

[!TIP]

Use the Reference Spec: The agent can use the references/gcp_pipeline_resource_provisioning_spec.md file as a template. It includes sample definitions for select supported resource types. Copy and adapt the required resource blocks into the deployment.yaml. Use gcloud beta orchestration-pipelines resource-types list when needed.

[!IMPORTANT]

Handling Secrets & Privacy (CRITICAL): NEVER hardcode plain-text secrets in deployment.yaml.

  • Sensitive Data (Secrets): Sensitive information such as passwords, API keys, and other sensitive information MUST be stored in Secret Manager and declared in the secrets: block of deployment.yaml.
  • Non-Sensitive Data (Variables): General configuration (e.g., dataset names, table IDs, regions) could be declared in the variables: block.
  • Substitution via {{ VAR }}: Both variables: and secrets: MUST be used as {{ VARIABLE_NAME }} substitutions in resource definitions.
  • No Creation: The agent MUST NOT use the framework to create new secrets. If gcloud indicates the secret does not exist, the agent MUST ask the user to create it manually and then re-verify.
  • Reference Only Policy: The agent's role is strictly limited to referencing existing secrets. The agent MUST NEVER read, print, or inspect the values of secrets.
  • Safe Deployment: The actual value injection happens during deployment execution. The agent only provides the reference.
  • Manual Secret Management: Advise the user to manage secret payloads and versions manually.

Step 4: Validation

The agent MUST validate the deployment.yaml before generating the deployment script. This ensures the configuration is syntactically correct and all variables are resolvable.

gcloud beta orchestration-pipelines validate --environment=<ENV_NAME>

Step 5: Deployment

Run the following command to deploy the resources to the target environment.

gcloud beta orchestration-pipelines deploy --environment=<ENV_NAME> --local

[!NOTE]

If a new transfer is being created, make sure to NOT remove the DTS transfer resource from deployment.yaml after it completes the run.

Definition of Done

  • deployment.yaml exists in the repository root with actual discovered values (no placeholders) and correct resource definitions.
  • The agent runs the deployment command to perform the deployment, and it executes successfully (exit code 0).

Source and attribution

Source:gemini-cli-extensions/data-agent-kit-starter-packinskills/gcp-pipeline-resource-provisioningat commit2df10e2

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from gemini-cli-extensions/data-agent-kit-starter-pack

Schema Mapping

gemini-cli-extensions

Plans source-to-target schema mappings for ETL, ELT, or data integration work, producing a documented Mapping Manifesto.

Data & Analytics215updated today

Resolving Mcp Region Configs

gemini-cli-extensions

Fixes unreplaced region placeholders in regional Google Cloud MCP server configs so missing MCP tools register.

DevOps & Cloud215updated today

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

Awaiting classification215updated today

Ml Best Practices

gemini-cli-extensions

Guides machine learning notebooks with step-by-step plans for clustering, forecasting, classification, regression and model comparison.

Data & Analytics215updated today

Managing Python Dependencies

gemini-cli-extensions

Guides agents to detect a Python project's dependency manager and install packages correctly instead of using global pip.

Software Development215updated today

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

Awaiting classification215updated today