Gcp Managed Airflow Dag Authoring

作者 gemini-cli-extensions2df10e25bbf7Apache-2.0215 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫今天更新

Guides the authoring and validation of Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA; formerly Cloud Composer). Covers environment context discovery, Airflow 2 vs 3 compatibility, authoring best practices, and local/remote validation processes. Use when creating or extending an Airflow DAG. Don't use when authoring Python code unrelated to Airflow DAGs.

AI 產生的概覽

指導為 Google Cloud Managed Service for Apache Airflow 撰寫與驗證 Apache Airflow DAG。

功能
此技能說明如何為 Managed Service for Apache Airflow(MSAA,前身為 Cloud Composer)撰寫 Apache Airflow DAG。內容涵蓋使用 gcloud 探索環境脈絡、套用冪等性與避免頂層程式碼等 Airflow 最佳實務、處理 Airflow 2 與 3 的相容性,以及在本機或目標環境中驗證 DAG。產出為 DAG 程式碼,以及驗證步驟與完成標準。
適用情境
適用於為 Managed Service for Apache Airflow 環境建立或擴充 Airflow DAG 的情境。不適用於與 Airflow DAG 無關的 Python 開發工作。
執行需求
僅為說明性內容,未附帶指令碼。內容提及 gcloud、composer-dev CLI、具備 airflow 的本機 Python 環境,以及 ruff 或 pylint 等檢查工具;目標環境驗證需要 GCP 存取權限,以及部署至 GCS 儲存桶的授權。

GCP Managed Airflow DAG Authoring Guide

This skill guides you through authoring and validating Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA; formerly Cloud Composer) environments.


Phase 1: Context Discovery

[!IMPORTANT] Before writing any DAG code, you MUST understand the constraints (e.g. version of Airflow) and capabilities of your target environment if user is willing to provide them.

1.1 Identify Target Environment & Access

Determine if you have direct access to the target Managed Airflow environment, local development environment or if you are working offline (only changing local files without validation).

  • If environment access is available: Use gcloud to inspect the environment (see Section 1.3).
  • If offline: Rely on user provided details.

1.2 Identify Development Environment

Determine if a local development environment is available.

  • Check if composer-dev CLI is installed.
  • Check if a local Python environment with airflow is available.

1.3 Inspect Target Environment (if available and requested)

Run the following commands to discover version constraints:

  1. Get Airflow/Image Version:

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.softwareConfig.imageVersion)"
  2. Get Installed Packages (Versions):

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.softwareConfig.pypiPackages)"
  3. Get DAGs GCS Bucket:

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.dagGcsPrefix)"

Phase 2: DAG Authoring Best Practices

2.1 General Airflow Best Practices

  • Idempotency: Every task SHOULD be idempotent. Running it multiple times with the same inputs (e.g., execution date) SHOULD produce the same result and not duplicate data.
  • No Top-Level Code Execution: Do NOT execute database queries, external API calls, or heavy computations at the top level of the DAG file (outside of tasks/operators). This code runs every few seconds during DAG parsing and will degrade performance.
  • Explicit Catchup: Always set catchup=False in the DAG definition unless historical backfilling is explicitly required.
  • Use Airflow Variables/Connections: Never hardcode credentials or environment-specific configs. Use Variable.get() (with deserialize_json=True if applicable) and BaseHook.get_connection(). Access variables via Jinja templates (e.g., {{ var.value.my_var }}) to avoid database calls during DAG parsing.

2.2 Airflow 2 vs Airflow 3 Compatibility

Reference @skill:gcp-managed-airflow-migrations to navigate adjusting the code to specific target Airflow version.


Phase 3: Validation Process

[!IMPORTANT] You MUST validate DAGs before concluding your task.

3.1 Local Validation (Offline/Pre-deployment)

3.1.1 Static Analysis & Linting

Use ruff or pylint if available.

bash
ruff check path/to/dag.py
  • If targeting Airflow 3, check with Airflow 3 rules if rulesets are available.
3.1.2 Local Dev Environment (composer-dev)

If the user has composer-dev configured:

  1. Copy the DAG to the local directory with DAGs:

    bash
    cp path/to/dag.py $(composer-dev describe <LOCAL_ENV> --format="value(dags_directory)")
  2. Verify parsing:

    bash
    composer-dev run-airflow-cmd <LOCAL_ENV> dags list-import-errors

3.2: Target Environment Validation

Only perform these steps if you have GCP access and are authorized to deploy to a target environment.

3.2.1 Deploy to GCS

Upload the DAG to the target environment's GCS bucket:

bash
gcloud storage cp path/to/dag.py gs://<TARGET_BUCKET>/dags/

3.2.2 Verify via Airflow CLI

Wait 1-2 minutes for the scheduler to parse the file, then run:

  1. Check for Import Errors:

    bash
    gcloud composer environments run <ENV_NAME> \    --location <REGION> \    dags list-import-errors

Pass Criteria: Output should be "No data found" or empty.

  1. Verify DAG is Listed:

    bash
    gcloud composer environments run <ENV_NAME> \    --location <REGION> \    dags list | grep <DAG_ID>

3.2.3 Monitor Cloud Logging

Check for runtime parsing errors in Cloud Logging:

query
resource.type="cloud_composer_environment"resource.labels.environment_name="<ENV_NAME>"log_id("airflow-scheduler")severity>=ERRORtextPayload:"<DAG_FILE_NAME>"

Definition of Done

  • DAG code adheres to Airflow version constraints of the target environment.
  • DAG code follows best practices (no top-level execution, idempotent if possible).
  • DAG parses locally without import errors.
  • (If environment is available) DAG is deployed to the target environment and verified to have no import errors.

來源與署名

來源:gemini-cli-extensions/data-agent-kit-starter-pack位於skills/gcp-managed-airflow-dag-authoring提交2df10e2

授權條款: Apache-2.0

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架

更多來自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

為 ETL、ELT 或資料整合任務規劃來源到目標的結構描述對應,產出文件化的對應宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修復區域性 Google Cloud MCP 伺服器設定中未取代的區域佔位符,讓缺少的 MCP 工具得以註冊。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分類215今天更新

Ml Best Practices

gemini-cli-extensions

為機器學習筆記本提供逐步方案,涵蓋分群、預測、分類、迴歸與模型比較。

Data & Analytics215今天更新

Managing Python Dependencies

gemini-cli-extensions

指導代理偵測 Python 專案的相依性管理器並正確安裝套件,而不是使用全域 pip。

Software Development215今天更新

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

待分類215今天更新