Gcp Managed Airflow Dag Authoring

作者 gemini-cli-extensions2df10e25bbf7Apache-2.0215 个星标收录于 2026年10月8日更新于 2026年10月8日仓库今天更新

Guides the authoring and validation of Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA; formerly Cloud Composer). Covers environment context discovery, Airflow 2 vs 3 compatibility, authoring best practices, and local/remote validation processes. Use when creating or extending an Airflow DAG. Don't use when authoring Python code unrelated to Airflow DAGs.

AI 生成的概览

指导为 Google Cloud Managed Service for Apache Airflow 编写和验证 Apache Airflow DAG。

功能
该技能指导如何为 Managed Service for Apache Airflow(MSAA,原 Cloud Composer)编写 Apache Airflow DAG。内容涵盖使用 gcloud 发现环境上下文、应用幂等性和避免顶层代码等 Airflow 最佳实践、处理 Airflow 2 与 3 的兼容性,以及在本地或目标环境中验证 DAG。产出为 DAG 代码以及验证步骤和完成标准。
适用场景
适用于为 Managed Service for Apache Airflow 环境创建或扩展 Airflow DAG 的场景。不适用于与 Airflow DAG 无关的 Python 开发工作。
运行要求
仅为说明性内容,不附带脚本。其中提到 gcloud、composer-dev CLI、带 airflow 的本地 Python 环境以及 ruff 或 pylint 等检查工具;目标环境验证需要 GCP 访问权限以及向 GCS 存储桶部署的授权。

GCP Managed Airflow DAG Authoring Guide

This skill guides you through authoring and validating Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA; formerly Cloud Composer) environments.


Phase 1: Context Discovery

[!IMPORTANT] Before writing any DAG code, you MUST understand the constraints (e.g. version of Airflow) and capabilities of your target environment if user is willing to provide them.

1.1 Identify Target Environment & Access

Determine if you have direct access to the target Managed Airflow environment, local development environment or if you are working offline (only changing local files without validation).

  • If environment access is available: Use gcloud to inspect the environment (see Section 1.3).
  • If offline: Rely on user provided details.

1.2 Identify Development Environment

Determine if a local development environment is available.

  • Check if composer-dev CLI is installed.
  • Check if a local Python environment with airflow is available.

1.3 Inspect Target Environment (if available and requested)

Run the following commands to discover version constraints:

  1. Get Airflow/Image Version:

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.softwareConfig.imageVersion)"
  2. Get Installed Packages (Versions):

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.softwareConfig.pypiPackages)"
  3. Get DAGs GCS Bucket:

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.dagGcsPrefix)"

Phase 2: DAG Authoring Best Practices

2.1 General Airflow Best Practices

  • Idempotency: Every task SHOULD be idempotent. Running it multiple times with the same inputs (e.g., execution date) SHOULD produce the same result and not duplicate data.
  • No Top-Level Code Execution: Do NOT execute database queries, external API calls, or heavy computations at the top level of the DAG file (outside of tasks/operators). This code runs every few seconds during DAG parsing and will degrade performance.
  • Explicit Catchup: Always set catchup=False in the DAG definition unless historical backfilling is explicitly required.
  • Use Airflow Variables/Connections: Never hardcode credentials or environment-specific configs. Use Variable.get() (with deserialize_json=True if applicable) and BaseHook.get_connection(). Access variables via Jinja templates (e.g., {{ var.value.my_var }}) to avoid database calls during DAG parsing.

2.2 Airflow 2 vs Airflow 3 Compatibility

Reference @skill:gcp-managed-airflow-migrations to navigate adjusting the code to specific target Airflow version.


Phase 3: Validation Process

[!IMPORTANT] You MUST validate DAGs before concluding your task.

3.1 Local Validation (Offline/Pre-deployment)

3.1.1 Static Analysis & Linting

Use ruff or pylint if available.

bash
ruff check path/to/dag.py
  • If targeting Airflow 3, check with Airflow 3 rules if rulesets are available.
3.1.2 Local Dev Environment (composer-dev)

If the user has composer-dev configured:

  1. Copy the DAG to the local directory with DAGs:

    bash
    cp path/to/dag.py $(composer-dev describe <LOCAL_ENV> --format="value(dags_directory)")
  2. Verify parsing:

    bash
    composer-dev run-airflow-cmd <LOCAL_ENV> dags list-import-errors

3.2: Target Environment Validation

Only perform these steps if you have GCP access and are authorized to deploy to a target environment.

3.2.1 Deploy to GCS

Upload the DAG to the target environment's GCS bucket:

bash
gcloud storage cp path/to/dag.py gs://<TARGET_BUCKET>/dags/

3.2.2 Verify via Airflow CLI

Wait 1-2 minutes for the scheduler to parse the file, then run:

  1. Check for Import Errors:

    bash
    gcloud composer environments run <ENV_NAME> \    --location <REGION> \    dags list-import-errors

Pass Criteria: Output should be "No data found" or empty.

  1. Verify DAG is Listed:

    bash
    gcloud composer environments run <ENV_NAME> \    --location <REGION> \    dags list | grep <DAG_ID>

3.2.3 Monitor Cloud Logging

Check for runtime parsing errors in Cloud Logging:

query
resource.type="cloud_composer_environment"resource.labels.environment_name="<ENV_NAME>"log_id("airflow-scheduler")severity>=ERRORtextPayload:"<DAG_FILE_NAME>"

Definition of Done

  • DAG code adheres to Airflow version constraints of the target environment.
  • DAG code follows best practices (no top-level execution, idempotent if possible).
  • DAG parses locally without import errors.
  • (If environment is available) DAG is deployed to the target environment and verified to have no import errors.

来源与署名

来源:gemini-cli-extensions/data-agent-kit-starter-pack位于skills/gcp-managed-airflow-dag-authoring提交2df10e2

许可证: Apache-2.0

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架

更多来自 gemini-cli-extensions/data-agent-kit-starter-pack 的技能

Schema Mapping

gemini-cli-extensions

为 ETL、ELT 或数据集成任务规划源到目标的模式映射,产出文档化的映射宣言。

Data & Analytics215今天更新

Resolving Mcp Region Configs

gemini-cli-extensions

修复区域级 Google Cloud MCP 服务器配置中未替换的区域占位符,使缺失的 MCP 工具得以注册。

DevOps & Cloud215今天更新

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

待分类215今天更新

Ml Best Practices

gemini-cli-extensions

为机器学习笔记本提供分步方案,涵盖聚类、预测、分类、回归和模型比较。

Data & Analytics215今天更新

Managing Python Dependencies

gemini-cli-extensions

指导代理检测 Python 项目的依赖管理器并正确安装依赖,而不是使用全局 pip。

Software Development215今天更新

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

待分类215今天更新