Skip to main content
A workflow is a program made of steps that run in an order you define. Each step is a Python function that performs a unit of work, and each step declares which step runs after it. Because the order comes from these connections instead of from top to bottom, steps that don’t depend on each other can run at the same time. This structure is called a directed acyclic graph (DAG); its steps connect in one direction and never loop back. The platform tracks every run’s inputs, outputs, and code automatically. Workflows on Anaconda Platform are built with Metaflow, a Python framework for data and ML pipelines.

How a workflow is structured

In code, a workflow is called a flow. A flow is a Python class based on FlowSpec, and the work happens in the class’s functions. The @step decorator marks a function as a step, and inside each step, self.next() declares which step runs next. Every flow has a start step and an end step. By convention, they’re named start and end:
Steps can also be marked explicitly with @step(start=True) and @step(end=True), which allows any step name:

What makes workflows different from scripts

A workflow is not just a Python script that runs top to bottom. The platform gives workflows several properties that plain scripts lack:
  • Each step is isolated. Steps can run on different machines with different resources. A preprocessing step might run on a small CPU instance while a training step runs on a GPU node.
  • Data flows between steps automatically. Values you assign to self in one step are available in all downstream steps, even across machines. The platform serializes, stores, and retrieves them transparently.
  • Every run is versioned. The platform records every run’s code, data, and results. You can inspect, compare, and reproduce past runs without manual bookkeeping.
  • Failures are recoverable. If a step fails, you can resume the flow from the point of failure without re-running the steps that already succeeded.

Workflows vs. deployments

Workflows and deployments are the two kinds of work you run on the platform:
  • A workflow runs to completion. It starts, executes its steps, and finishes. Use workflows for training models, processing data, running evaluations, and any work with a defined end.
  • A deployment stays running. It serves requests continuously until you stop it. Use deployments for model endpoints, APIs, dashboards, and services that need to be available on demand.
Both run on the compute assigned to their perimeter.

Building workflows

To build and run your first flow, see Connect to Anaconda Platform and run your first flow. For flow-building features like parameters, branching, and parallelism, see the Metaflow documentation.