Gcp Pipeline Resource Provisioning

作者 gemini-cli-extensions2df10e25bbf7Apache-2.0215 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Automates declarative resource creation and provisioning for data pipelines, supporting BigQuery, Dataform, Dataproc, BigQuery Data Transfer Service (DTS), and other resources. It manages environment-specific configurations (dev, staging, prod) through a deployment.yaml file. Use when: - Modifying or creating deployment.yaml for deployment settings. - Resolving environment-specific variables (e.g., Project IDs, Regions) for deployment. - Provisioning supported infrastructure like BigQuery datasets/tables, Dataform resources, or DTS resources via deployment.yaml. Do not use when: - Resources already exist. - Managing resources not supported by `gcloud beta orchestration-pipelines resource-types list`. - Managing general cloud infrastructure (VMs, networks, Kubernetes, IAM policies), which are better suited for Terraform. - Infrastructure spans multiple cloud providers (AWS, Azure, etc.). - Already uses Terraform for the target resources.

仅含说明DevOps & Cloud
AI 生成的概览

生成并部署 deployment.yaml,用于跨环境预置 GCP 数据管道资源。

功能
该技能引导代理在仓库根目录创建或更新 deployment.yaml,用于声明 BigQuery 数据集和表、Dataform 资源、Dataproc 以及 BigQuery 数据传输服务等 GCP 数据管道资源。它涵盖使用 gcloud 和 git 命令发现环境参数、为 dev、stage 和 prod 解析环境专属取值、引用 Secret Manager 中的密钥、校验配置并执行部署命令。其产出是 deployment.yaml 文件以及一次已执行的部署。
适用场景
适用于通过 deployment.yaml 预置或修改受支持的 GCP 管道资源,或解析项目 ID、区域等环境专属变量。不适用于已存在的资源、虚拟机、网络、Kubernetes 或 IAM 等通用云基础设施、多云环境,或已由 Terraform 管理的目标。
运行要求
需要带 beta orchestration-pipelines 命令的 gcloud CLI、已认证的 GCP 项目、用于获取提交 SHA 的 git,以及访问 Secret Manager 以引用密钥。该技能不附带脚本,依赖参考规范文件和用户提供的环境名称。

How to use this skill

Create or update existing deployment.yaml file and deploy resources. All configuration files MUST be maintained together in the repository root.

Mandatory labels

[!IMPORTANT]

Whenever you generate resource definitions in deployment.yaml, you MUST directly populate the datacloud label under definition.labels for every resource to track the source of creation. Determine the value based on your current IDE environment:

  • For Antigravity, set datacloud: "antigravity"
  • For VS Code, set datacloud: "vscode"
  • For any other environment, set datacloud: "other"

Do not use a variable substitution for this label; hardcode the appropriate string value directly into each resource definition (e.g., replacing __REQUIRED_LABEL__ placeholders).

Special rule for BigQuery DTS Ingestion: Whenever you generate a bigquerydatatransfer.transferConfig in deployment.yaml, you MUST also explicitly define its target destination bigquery.dataset in the same file and apply the datacloud label to it. You must do this even if the dataset already exists, to ensure the destination dataset's labels are patched and updated.

Step 1: Supported Resource Types

The framework supports deploying various GCP resources. To see the comprehensive list of supported resource types, run the following command:

gcloud beta orchestration-pipelines resource-types list

Refer to: references/gcp-pipeline-resource-provisioning_spec.md to understand the template for deployment.yaml.

Step 2: Discover Environment Parameters

Before generating configurations, discover the actual values for the target project, region, environment, and commit SHA.

[!TIP]

If deployment.yaml already exists in the repository root, prioritize extracting project and region from the target environment configuration (e.g., dev).

  1. Project ID:

    bash
    gcloud config get project
  2. Project Number:

    bash
    gcloud projects describe $(gcloud config get project) --format="value(projectNumber)"
  3. Region:

    bash
    gcloud config get-value compute/region
  4. Commit SHA:

    bash
    git rev-parse HEAD
  5. Environment Name: If initialization is needed, you MUST ask the user for the environment name. If the user does not provide it, use dev as the default.

[!TIP]

Use these commands to replace placeholders like YOUR_PROJECT_ID with actual values. Always remove associated comments that start with TODO once replaced.

Step 3: Generate or update deployment.yaml

Create or update deployment.yaml in the repository root. This file maps supported environments (dev, stage, prod) to their specific configurations and resources.

[!TIP]

Use the Reference Spec: The agent can use the references/gcp_pipeline_resource_provisioning_spec.md file as a template. It includes sample definitions for select supported resource types. Copy and adapt the required resource blocks into the deployment.yaml. Use gcloud beta orchestration-pipelines resource-types list when needed.

[!IMPORTANT]

Handling Secrets & Privacy (CRITICAL): NEVER hardcode plain-text secrets in deployment.yaml.

  • Sensitive Data (Secrets): Sensitive information such as passwords, API keys, and other sensitive information MUST be stored in Secret Manager and declared in the secrets: block of deployment.yaml.
  • Non-Sensitive Data (Variables): General configuration (e.g., dataset names, table IDs, regions) could be declared in the variables: block.
  • Substitution via {{ VAR }}: Both variables: and secrets: MUST be used as {{ VARIABLE_NAME }} substitutions in resource definitions.
  • No Creation: The agent MUST NOT use the framework to create new secrets. If gcloud indicates the secret does not exist, the agent MUST ask the user to create it manually and then re-verify.
  • Reference Only Policy: The agent's role is strictly limited to referencing existing secrets. The agent MUST NEVER read, print, or inspect the values of secrets.
  • Safe Deployment: The actual value injection happens during deployment execution. The agent only provides the reference.
  • Manual Secret Management: Advise the user to manage secret payloads and versions manually.

Step 4: Validation

The agent MUST validate the deployment.yaml before generating the deployment script. This ensures the configuration is syntactically correct and all variables are resolvable.

gcloud beta orchestration-pipelines validate --environment=<ENV_NAME>

Step 5: Deployment

Run the following command to deploy the resources to the target environment.

gcloud beta orchestration-pipelines deploy --environment=<ENV_NAME> --local

[!NOTE]

If a new transfer is being created, make sure to NOT remove the DTS transfer resource from deployment.yaml after it completes the run.

Definition of Done

  • deployment.yaml exists in the repository root with actual discovered values (no placeholders) and correct resource definitions.
  • The agent runs the deployment command to perform the deployment, and it executes successfully (exit code 0).

来源与署名

来源:gemini-cli-extensions/data-agent-kit-starter-pack位于skills/gcp-pipeline-resource-provisioning提交2df10e2

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

为 ETL、ELT 或数据集成任务规划源到目标的模式映射,产出文档化的映射宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修复区域级 Google Cloud MCP 服务器配置中未替换的区域占位符,使缺失的 MCP 工具得以注册。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分类215今天更新

Ml Best Practices

gemini-cli-extensions

为机器学习笔记本提供分步方案,涵盖聚类、预测、分类、回归和模型比较。

Data & Analytics215今天更新

Managing Python Dependencies

gemini-cli-extensions

指导代理检测 Python 项目的依赖管理器并正确安装依赖,而不是使用全局 pip。

Software Development215今天更新

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

待分类215今天更新