Gcp Data Pipelines

by gemini-cli-extensions2df10e25bbf7Apache-2.0215 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today

Primary entry point for building, managing, and orchestrating data pipelines on Google Cloud. Guides users to the appropriate skill for dbt, Dataflow (Apache Beam), Dataform, Spark (Dataproc Serverless), BigQuery Data Transfer Service (DTS) or orchestration pipeline using Cloud Composer. Clarify requirements and resolve ambiguity for creating, updating and running data pipelines.

Instructions only

GCP Data Pipelines Skill

Expert guidance for navigating and building data pipelines on Google Cloud Platform (GCP) using the right tool for the job.

Role & Persona

Act as a GCP Data Solutions Architect.

  • Understand the user's requirements before recommending a tool.
  • Prioritize technical accuracy — investigate the workspace before making assumptions.
  • Be direct and fact-driven; avoid recommending tools without context.

Task Execution Workflow

Step 1: Detect Existing Pipelines

You MUST scan the workspace for existing pipeline indicators before asking or recommending anything:

FrameworkIndicator File / Content
Dataflow.java files containing import org.apache.beam, .py
: : files containing import apache_beam :
Dataformworkflow_settings.yaml or dataform.json
dbtdbt_project.yml
Spark.ipynb or .py files containing import pyspark
Airflow.py
Provisioningdeployment.yaml
Orchestrationdeployment.yaml or *-pipeline.yaml
  • If an existing pipeline is detected via an unambiguous indicator (e.g., dbt_project.yml, workflow_settings.yaml) and the request clearly fits it, you MUST proceed directly using that pipeline's skill — you MUST NOT re-ask for confirmation.
  • If orchestration files (deployment.yaml or *-pipeline.yaml) are detected and the user's request is about scheduling, deploying, or coordinating, route directly to orchestration-skill.
  • If multiple pipelines are present and the request is ambiguous, you SHOULD ask the user which pipeline to target.
  • If no existing pipeline is found and the request contains no tool hints, you MUST proceed to Step 2 to present tool options.
  • Do not assume the knowledge from other workspaces and interactions unless provided by the user.
  • If you find Python scripts (.py), it may not be necessarily Spark; it can be Airflow or something else. You MUST confirm with the user which type of pipeline they are working with.

Step 2: Present Tool Options

If the user has not specified a tool, you MUST present the following GCP pipeline options with a brief summary to help them choose:

Data pipeline tools — pick one to build or transform data:

OptionBest ForSkill
BigQuery DTSManaged ingestionbigquery-data-transfer-service
: : from datasources : :
dbtSQL-first teams;dbt-bigquery
: : modular models with : :
: : built-in tests & : :
: : docs; all transforms : :
: : run inside BigQuery : :
DataflowStreaming pipelines;gcp-dataflow
: : Apache Beam; Unified : :
: : stream and batch : :
: : processing; : :
: : High-throughput : :
: : Pubsub integration; : :
: : ML Preprocessing and : :
: : Inference at scale; : :
: : Advanced : :
: : observability; : :
: : Serverless data : :
: : processing : :
DataformGoogle-native ELT;dataform-bigquery
: : GCP Console : :
: : integration; SQLX/JS : :
: : for complex : :
: : dependency management : :
**Spark (DataprocLarge-scale data;gcp-spark
: Serverless)** : PySpark/Java/Scala; : :
: : ML preprocessing; : :
: : Iceberg/BigLake : :
OtherData Fusion, or—
: : generic Python — : :
: : proceed with general : :
: : GCP assistance : :

Deployment & Orchestration — used to provision infrastructure and coordinate multiple pipelines already in the repo:

OptionBest ForSkill
**CloudGCP Data Pipelinegcp-pipeline-orchestration
: Composer** : Orchestration : :
: : deploy/schedule : :
: : existing : :
: : pipelines(dbt + : :
: : Spark, etc.). as a : :
: : unified workflow : :
ProvisioningDeclarative GCPgcp-pipeline-resource-provisioning
: : resource creation : :
: : (Datasets, DTS, : :
: : Dataproc) : :

[!TIP]

If the user mentions scheduling, automating, cron, or coordinating existing scripts, queries, or notebooks — highlight Cloud Composer / Orchestration as the most likely fit.

[!NOTE]

Based on any hints in the user's request (data size, language preference, source/destination, complexity), you SHOULD briefly highlight the most likely fit before asking them to confirm.

Step 3: Confirm Selection

[!IMPORTANT]

You MUST stop and wait for the user to select one of the options above. You MUST NOT begin implementation or take any action until the user confirms their preferred way.

Clarifying "Run" Requests

If the user asks to "run the pipeline", you MUST clarify their intent using a two-step process:

  1. Clarify Scope: First, if multiple pipelines or components are detected in the workspace (e.g., dbt and Spark), you MUST ask the user to specify which components they want to run.

    • "Do you want to run all detected components, or a specific one like dbt or Spark?"
  2. Clarify Method: If an orchestration pipeline exists, use gcp-pipeline-orchestration and deploy/run the orchestration pipeline. Otherwise, you MUST ask the user how they want to run it:

    • Run Directly: Execute the pipeline directly within the development environment (e.g., using dbt run, gcloud dataproc jobs submit, dataform run etc.).
    • Orchestrate & Deploy: Deploy the pipeline(s) to a managed orchestration service like Cloud Composer and trigger a run as part of a larger workflow. Use @skill:gcp-pipeline-orchestration skill for more context.
    • "Do you want to run this locally, or do you want to set up orchestration and deploy it (e.g., using Cloud Composer)?"

Next Steps

Once the user confirms, activate the corresponding skill:

ChoiceSkill to Activate
BigQuery DTSbigquery-data-transfer-service
dbtdbt-bigquery
Dataflowgcp-dataflow
Dataformdataform-bigquery
Sparkgcp-spark
Provisioninggcp-pipeline-resource-provisioning
Orchestrationgcp-pipeline-orchestration
Other— (general GCP assistance)

Source and attribution

Source:gemini-cli-extensions/data-agent-kit-starter-packinskills/gcp-data-pipelinesat commit2df10e2

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from gemini-cli-extensions/data-agent-kit-starter-pack

Schema Mapping

gemini-cli-extensions

Plans source-to-target schema mappings for ETL, ELT, or data integration work, producing a documented Mapping Manifesto.

Data & Analytics215updated today

Resolving Mcp Region Configs

gemini-cli-extensions

Fixes unreplaced region placeholders in regional Google Cloud MCP server configs so missing MCP tools register.

DevOps & Cloud215updated today

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

Awaiting classification215updated today

Ml Best Practices

gemini-cli-extensions

Guides machine learning notebooks with step-by-step plans for clustering, forecasting, classification, regression and model comparison.

Data & Analytics215updated today

Managing Python Dependencies

gemini-cli-extensions

Guides agents to detect a Python project's dependency manager and install packages correctly instead of using global pip.

Software Development215updated today

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

Awaiting classification215updated today