
Gcp Dataflow
by gemini-cli-extensions2df10e25bbf7Apache-2.0215 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated today
Guides writing, packaging, executing, and troubleshooting Apache Beam pipelines on Dataflow. Use when creating new pipelines, configuring Flex Templates, or analyzing performance of Dataflow jobs. Capabilities include Java/Python/Go setup, Cloud Build integration, and deep diagnostic analysis of job health and autoscaling. Use when: - Creating an Apache Beam Dataflow pipeline. - Creating a Google Dataflow Flex Template. - Using an existing Google Dataflow Template. - Debugging Dataflow pipeline - Troubleshooting Dataflow pipeline - Analyzing Performance of Dataflow pipeline. Key capabilities: Java/Python/Go project setup, Flex Templates (with Cloud Build), and diagnostics for streaming job health, bottlenecks, and autoscaling. Do NOT use for: - General GCP resource management unrelated to Dataflow. - Issues with other GCP services (e.g., GCE, GCS, BigQuery) unless directly impacting Dataflow pipeline execution. - Pipeline technologies other than Apache Beam on Dataflow.
- 2df10e25bbf7Currentcommit 2df10e2Published Oct 8, 2026
Source and attribution
Source:gemini-cli-extensions/data-agent-kit-starter-packinskills/gcp-dataflowat commit2df10e2
License: Apache-2.0
Content belongs to its original authors. SourceWeft indexes it from a public repository.
More from gemini-cli-extensions/data-agent-kit-starter-pack

Schema Mapping
gemini-cli-extensions
Plans source-to-target schema mappings for ETL, ELT, or data integration work, producing a documented Mapping Manifesto.

Resolving Mcp Region Configs
gemini-cli-extensions
Fixes unreplaced region placeholders in regional Google Cloud MCP server configs so missing MCP tools register.

Notebook Guidance
gemini-cli-extensions
This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

Ml Best Practices
gemini-cli-extensions
Guides machine learning notebooks with step-by-step plans for clustering, forecasting, classification, regression and model comparison.

Managing Python Dependencies
gemini-cli-extensions
Guides agents to detect a Python project's dependency manager and install packages correctly instead of using global pip.

Google Cloud Storage Fuse
gemini-cli-extensions
Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).
More in DevOps & Cloud

M5 Onboard
anthropics
Provisions M5Stack ESP32 boards by detecting them on USB, flashing UIFlow 2.0 firmware, and installing a MicroPython app bundle.

Runbook
anthropics
Creates or updates step-by-step operational runbooks for recurring tasks, including troubleshooting, rollback and escalation.

Incident Response
anthropics
Guides an incident response workflow: severity triage, status updates, mitigation tracking, and blameless postmortems.

Deploy Checklist
anthropics
Generates a pre-deployment readiness checklist covering pre-deploy, deploy, post-deploy and rollback triggers.

Spanner Basics
Guides Google Cloud Spanner administration, schema design, querying and performance diagnosis.

Secops Cases
Manages Google Security Operations SOAR cases across their lifecycle via MCP tools.