Domino Flows

by dominodatalabd86698d74d56No license7 starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 2 days ago

Orchestrate multi-step ML workflows using Domino Flows (built on Flyte). Define DAGs with typed inputs/outputs, heterogeneous environments, automatic lineage, and reproducibility. Use when building data pipelines, multi-stage training workflows, or processes requiring orchestration and monitoring.

Instructions onlyDevOps & CloudAI & Agents
AI-generated overview

Guides building multi-step ML workflows as Domino Flows DAGs on Flyte, with typed tasks and stage scripts.

What it does
This skill supplies reference knowledge for orchestrating multi-step machine learning workflows with Domino Flows, which is built on Flyte. It explains DAG concepts, typed inputs and outputs, per-task environments, lineage and reproducibility, and shows how to define tasks with DominoJobTask and DominoJobConfig inside a workflow. It also documents the stage-script pattern that reads from workflow inputs and writes to workflow outputs, plus how to run a flow remotely. It produces guidance and code patterns rather than executable artifacts.
When to use it
Use it when building data pipelines, multi-stage training workflows, or ETL processes that need orchestration, monitoring, lineage and reproducibility. It is also relevant when tasks must run in different environments or be scheduled and triggered. It is not intended for single-step processes, real-time inference, or tasks sharing mutable state.
Requirements
Requires familiarity with Domino Flows and Flyte, the flytekit and flytekitplugins.domino.task Python packages, a Domino environment with remote job execution, and repository access for committing and pushing code. It ships no scripts; it is instructions and reference documents only.

Domino Flows Skill

This skill provides comprehensive knowledge for orchestrating ML workflows using Domino Flows, built on the Flyte platform.

Key Concepts

What are Domino Flows?

Domino Flows enable:

  • DAG-based orchestration: Define workflows as directed acyclic graphs
  • Typed interfaces: Strong typing for inputs and outputs
  • Heterogeneous environments: Different environments per task
  • Automatic lineage: Track data and model provenance
  • Reproducibility: Version-controlled workflows
  • Scalability: Distributed execution across compute resources

Core Components

ComponentDescription
TaskSingle unit of work (runs as a Domino Job)
WorkflowDAG connecting tasks
ArtifactTyped input/output passed between tasks
Launch PlanConfigured workflow execution

Related Documentation

  • FLOW-BASICS.md - DAG concepts, task definitions
  • EXAMPLES.md - Common flow patterns

Quick Start

⚠️ Critical: Domino Flows does NOT support native Flyte @task decorators. Tasks must use DominoJobTask + DominoJobConfig. Only @workflow is unchanged.

Basic Flow

Each task runs as a Domino Job. Stage scripts read from /workflow/inputs/<name> and write to /workflow/outputs/o0. Pass PYTHONPATH=/mnt/code in the command.

python
from flytekit import workflowfrom flytekitplugins.domino.task import DominoJobConfig, DominoJobTask
preprocess_task = DominoJobTask(    name="Preprocess Data",    domino_job_config=DominoJobConfig(        Command="bash -c 'PYTHONPATH=/mnt/code python /mnt/code/stages/preprocess.py'",    ),    inputs={"input_path": str},    outputs={"o0": str},    use_latest=True,)
train_task = DominoJobTask(    name="Train Model",    domino_job_config=DominoJobConfig(        Command="bash -c 'PYTHONPATH=/mnt/code python /mnt/code/stages/train.py'",    ),    inputs={"preprocess_output": str},    outputs={"o0": str},    use_latest=True,)
@workflowdef training_pipeline(input_path: str = "/mnt/data/raw.csv") -> str:    preprocess_output = preprocess_task(input_path=input_path)    result = train_task(preprocess_output=preprocess_output)    return result

Stage Script Pattern

python
# stages/preprocess.pyimport json, os
INPUTS, OUTPUTS = "/workflow/inputs", "/workflow/outputs"
def main():    input_path = open(f"{INPUTS}/input_path").read().strip()    # ... do work ...    os.makedirs(OUTPUTS, exist_ok=True)    with open(f"{OUTPUTS}/o0", "w") as f:        f.write(json.dumps({"output_path": "/mnt/artifacts/processed.parquet"}))
if __name__ == "__main__":    main()

Running the Flow

bash
# Always commit and push first — jobs run against remote repo stategit add -A && git commit -m "..." && git push
# Trigger remotelyPYTHONPATH=/mnt/code pyflyte run --remote \    my_flow.py training_pipeline \    --input_path "/mnt/data/raw.csv"

When to Use Flows

Good Use Cases

  • Data processing → Model training pipelines
  • ETL with ML steps
  • Multi-stage training with different environments
  • Processes requiring reproducibility and lineage
  • Scheduled/triggered workflows

Not Ideal For

  • Single dataset with many small computations
  • Tasks that write to mutable shared state
  • Simple single-step processes
  • Real-time inference (use Model APIs instead)

Documentation Links

Source and attribution

Source:dominodatalab/domino-claude-plugininskills/flowsat commitd86698d

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal