Gcp Managed Airflow Dag Authoring

by gemini-cli-extensions2df10e25bbf7Apache-2.0Listed Oct 8, 2026Updated Oct 8, 2026

Guides the authoring and validation of Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA; formerly Cloud Composer). Covers environment context discovery, Airflow 2 vs 3 compatibility, authoring best practices, and local/remote validation processes. Use when creating or extending an Airflow DAG. Don't use when authoring Python code unrelated to Airflow DAGs.

AI-generated overview

Guides authoring and validating Apache Airflow DAGs for Google Cloud Managed Service for Apache Airflow.

What it does
This skill walks through authoring Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA, formerly Cloud Composer). It covers discovering environment context with gcloud, applying Airflow best practices such as idempotency and avoiding top-level code, handling Airflow 2 versus 3 compatibility, and validating DAGs locally or against a target environment. It produces DAG code plus validation steps and a definition of done.
When to use it
Use it when creating or extending an Airflow DAG for a Managed Service for Apache Airflow environment. It is not intended for Python work unrelated to Airflow DAGs.
Requirements
Instructions only, with no bundled scripts. It references gcloud, the composer-dev CLI, a local Python environment with airflow, and linters such as ruff or pylint; target-environment validation needs GCP access and authorization to deploy to a GCS bucket.

GCP Managed Airflow DAG Authoring Guide

This skill guides you through authoring and validating Apache Airflow DAGs for Managed Service for Apache Airflow (MSAA; formerly Cloud Composer) environments.


Phase 1: Context Discovery

[!IMPORTANT] Before writing any DAG code, you MUST understand the constraints (e.g. version of Airflow) and capabilities of your target environment if user is willing to provide them.

1.1 Identify Target Environment & Access

Determine if you have direct access to the target Managed Airflow environment, local development environment or if you are working offline (only changing local files without validation).

  • If environment access is available: Use gcloud to inspect the environment (see Section 1.3).
  • If offline: Rely on user provided details.

1.2 Identify Development Environment

Determine if a local development environment is available.

  • Check if composer-dev CLI is installed.
  • Check if a local Python environment with airflow is available.

1.3 Inspect Target Environment (if available and requested)

Run the following commands to discover version constraints:

  1. Get Airflow/Image Version:

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.softwareConfig.imageVersion)"
  2. Get Installed Packages (Versions):

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.softwareConfig.pypiPackages)"
  3. Get DAGs GCS Bucket:

    bash
    gcloud composer environments describe <ENV_NAME> \    --location <REGION> \    --format="value(config.dagGcsPrefix)"

Phase 2: DAG Authoring Best Practices

2.1 General Airflow Best Practices

  • Idempotency: Every task SHOULD be idempotent. Running it multiple times with the same inputs (e.g., execution date) SHOULD produce the same result and not duplicate data.
  • No Top-Level Code Execution: Do NOT execute database queries, external API calls, or heavy computations at the top level of the DAG file (outside of tasks/operators). This code runs every few seconds during DAG parsing and will degrade performance.
  • Explicit Catchup: Always set catchup=False in the DAG definition unless historical backfilling is explicitly required.
  • Use Airflow Variables/Connections: Never hardcode credentials or environment-specific configs. Use Variable.get() (with deserialize_json=True if applicable) and BaseHook.get_connection(). Access variables via Jinja templates (e.g., {{ var.value.my_var }}) to avoid database calls during DAG parsing.

2.2 Airflow 2 vs Airflow 3 Compatibility

Reference @skill:gcp-managed-airflow-migrations to navigate adjusting the code to specific target Airflow version.


Phase 3: Validation Process

[!IMPORTANT] You MUST validate DAGs before concluding your task.

3.1 Local Validation (Offline/Pre-deployment)

3.1.1 Static Analysis & Linting

Use ruff or pylint if available.

bash
ruff check path/to/dag.py
  • If targeting Airflow 3, check with Airflow 3 rules if rulesets are available.
3.1.2 Local Dev Environment (composer-dev)

If the user has composer-dev configured:

  1. Copy the DAG to the local directory with DAGs:

    bash
    cp path/to/dag.py $(composer-dev describe <LOCAL_ENV> --format="value(dags_directory)")
  2. Verify parsing:

    bash
    composer-dev run-airflow-cmd <LOCAL_ENV> dags list-import-errors

3.2: Target Environment Validation

Only perform these steps if you have GCP access and are authorized to deploy to a target environment.

3.2.1 Deploy to GCS

Upload the DAG to the target environment's GCS bucket:

bash
gcloud storage cp path/to/dag.py gs://<TARGET_BUCKET>/dags/

3.2.2 Verify via Airflow CLI

Wait 1-2 minutes for the scheduler to parse the file, then run:

  1. Check for Import Errors:

    bash
    gcloud composer environments run <ENV_NAME> \    --location <REGION> \    dags list-import-errors

Pass Criteria: Output should be "No data found" or empty.

  1. Verify DAG is Listed:

    bash
    gcloud composer environments run <ENV_NAME> \    --location <REGION> \    dags list | grep <DAG_ID>

3.2.3 Monitor Cloud Logging

Check for runtime parsing errors in Cloud Logging:

query
resource.type="cloud_composer_environment"resource.labels.environment_name="<ENV_NAME>"log_id("airflow-scheduler")severity>=ERRORtextPayload:"<DAG_FILE_NAME>"

Definition of Done

  • DAG code adheres to Airflow version constraints of the target environment.
  • DAG code follows best practices (no top-level execution, idempotent if possible).
  • DAG parses locally without import errors.
  • (If environment is available) DAG is deployed to the target environment and verified to have no import errors.

Source and attribution

Source:gemini-cli-extensions/data-agent-kit-starter-packinskills/gcp-managed-airflow-dag-authoringat commit2df10e2

License: Apache-2.0

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal

More from gemini-cli-extensions/data-agent-kit-starter-pack

Schema Mapping

gemini-cli-extensions

Plans source-to-target schema mappings for ETL, ELT, or data integration work, producing a documented Mapping Manifesto.

Data & AnalyticsOct 8, 2026

Resolving Mcp Region Configs

gemini-cli-extensions

Fixes unreplaced region placeholders in regional Google Cloud MCP server configs so missing MCP tools register.

DevOps & CloudOct 8, 2026

Notebook Guidance

gemini-cli-extensions

This skill guides the use of Jupyter notebooks for data analysis, exploration, and visualization, particularly with BigQuery. It outlines best practices for notebook execution and validation (supporting both cell-by-cell execution and full notebook generation depending on tool availability), library installation, and structuring notebooks for clarity. It also covers specific rules for data cleaning, plotting, and integrating with BigQuery SQL and machine learning workflows. Relevant when any of the following conditions are true: 1. The user request involves a data analysis, data exploration, data visualization, or data insights task that requires multiple steps, queries, or visualizations to answer. 2. The user explicitly requests a notebook (.ipynb). 3. You are creating, editing, or executing cells in a Jupyter notebook. 4. You need to query BigQuery from within a notebook. DO NOT use the Python BigQuery client library; instead, you MUST use the `%%bqsql` magics explained in this skill.

Awaiting classificationOct 8, 2026

Ml Best Practices

gemini-cli-extensions

Guides machine learning notebooks with step-by-step plans for clustering, forecasting, classification, regression and model comparison.

Data & AnalyticsOct 8, 2026

Managing Python Dependencies

gemini-cli-extensions

Guides agents to detect a Python project's dependency manager and install packages correctly instead of using global pip.

Software DevelopmentOct 8, 2026

Google Cloud Storage Fuse

gemini-cli-extensions

Mounts Cloud Storage buckets as a POSIX file system with Cloud Storage FUSE (gcsfuse). Use when you need to interact with gcsfuse — decide whether FUSE, native gs:// reads, or Filestore/Managed Lustre fits a workload, deploy tuned mounts on GKE, Compute Engine, or Cloud Run, enable and size the file, stat, and list caches, tune mount flags or config-file settings, apply workload profiles, keep ML checkpointing safe (rename atomicity, hierarchical namespace, close-time finalization, concurrent writers), or diagnose slow training, low throughput, or GCS bill spikes on existing mounts with gcsfuse metrics. Covers mount semantics, the gcsfuse CLI and config file, the GKE gcsfuse CSI driver (Workload Identity principal:// bindings, profile StorageClasses, sidecar sizing), and Cloud Run volume mounts. Don't use for bucket administration or data management without a mount (google-cloud-storage-basics) or for fully POSIX-compliant shared file systems (Filestore, Managed Lustre).

Awaiting classificationOct 8, 2026