Skip to main content
After setting up an empty project, you can begin adding your own components. Fundamentally, all projects are composed of these top-level components:
The elements of a project: workflows, deployments, assets, and evaluations organized under branches

Flows

Flows refer to Metaflow flows, often interconnected through events. They form the backbone of your projects, handling data processing and ETL, model training and finetuning, autonomous and batch inference, and any other types of background processing and high-performance computing. In projects, flows are stored under a subdirectory flows, one Metaflow flow (named flow.py) per subdirectory, alongside any supporting Python modules and packages. As a best practice, it is useful to add a README.md file for each flow describing its role. They will be surfaced in the UI as well.

Writing a ProjectFlow

Importantly, project flows should subclass from ProjectFlow instead of Metaflow’s standard FlowSpec. In other words, simply author your flows like this:
This leverages Metaflow’s BaseFlow pattern to enrich flows with functionality related to the project structure. Besides this small detail, you can use all Metaflow features in your flows. A typical flow hierarchy in a project repository ends up looking like this:
flows
etl
flow.py
README.md
feature_transformations.py
sql
process_data.sql
train_model
flow.py
README.md
model.py

Deployments

Deployments are microservices that serve requests through real-time APIs. Common use cases:

Flows and deployments working together

The platform’s strength comes from the tight connection between flows and deployments, bridging the offline and online worlds:
  • A flow can continuously update a database for RAG, which a deployed agent then uses in real time.
  • A deployed app can monitor model performance and trigger a retraining flow.
  • A flow can deploy a model endpoint programmatically, for instance, whenever a new model has been trained.

Deployments in a project

In your project, place deployments in the deployments directory. Each deployment is defined by a configuration file, config.yml, as documented in Deployments deep dive. You can define dependencies for the deployment in a standard requirements.txt or pyproject.toml. As with flows, it is recommended to add a README.md for each deployment. The project hierarchy looks like this:
deployments
monitoring_dashboard
streamlit_app.py
config.yml
pyproject.toml
README.md
model_endpoint
fastapi_server.py
config.yml
pyproject.toml
README.md
support_agent
agent.py
config.yml
pyproject.toml
README.md

Deployment commands

When obproject-deploy deploys apps, it runs from the project root directory, so commands in your config.yml must use paths relative to the project root, not the deployment directory. For example, if your project structure is:
my_project
deployments
dashboard
app.py
config.yml
api
main.py
config.yml
src
shared_utils.py
Write the config.yml commands with paths relative to the project root:
Running from the project root has two benefits:
  • Shared imports: Apps can import modules from src/ directly (from my_module import ...).
  • Consistency: All paths are relative to the same root, making them predictable.
Use file paths for CLI tools (streamlit run deployments/dashboard/app.py) and Python module paths for WSGI/ASGI servers (gunicorn ... deployments.api.main:app).

Code

Effective management of software dependencies is essential for building production-quality projects and enabling rapid iteration and collaboration. A project has one platform configuration file, obproject.toml, and one or more dependency manifests, such as pyproject.toml or requirements.txt. The platform configuration says what the project is and how it deploys; the dependency manifests say what packages the code needs: A typical project consists of multiple layers of software dependencies:
  • Code defining flows and deployments, organized into subdirectories.
  • Project-level shared libraries under the src directory.
  • Organization-level libraries shared across projects.
  • Third-party dependencies, such as pandas and torch, declared at the flow, deployment, or project level.
For example, consider the following project that trains a fraud detection model and deploys it for real-time inference:
fraud_detection_model
obproject.toml
pyproject.toml
README.md
src
feature_encoders
__init__.py
feature_encoder.py
flows
trainer
flow.py
mymodel.py
README.md
deployments
inference
fastapi_server.py
config.yml
README.md

Code for flows and deployments

In addition to the entrypoint file (flow.py) or deployment server, each flow or deployment can include supporting modules and packages, such as mymodel.py in the tree above.

Project-level shared libraries

Place libraries shared within a project under the src directory, as packages. For example, a feature_encoders package used both during training and inference ensures offline-online consistency of features:
src
feature_encoders
__init__.py
feature_encoder.py
In each package’s __init__.py, include:
This ensures the package gets included in the Metaflow code package when deployed. When you run obproject-deploy, it automatically sets up PYTHONPATH so your flows and apps can import these modules directly, such as from feature_encoders import MyEncoder.

Organization-level libraries

Libraries shared across projects can be handled in two ways:
  • If you can set METAFLOW_PACKAGE_POLICY in packages, simply pip install them as usual or add them to your PYTHONPATH. Once you import them in your flows and deployments, they get packaged automatically. This is a convenient option for private packages, even if they are not pip install-able from a package repository.
  • If the shared libraries are pushed to a package repository, private or public, treat them like third-party dependencies.

Declaring dependencies

You can declare dependencies at three levels in a project:
  • Per flow or step: Use Metaflow’s @pypi or @conda decorators in your flow code.
  • Per deployment: Add a requirements.txt or pyproject.toml to the deployment directory, or declare dependencies in its config.yml.
  • Project-wide: Place a pyproject.toml at the root of the project next to obproject.toml. For example:
A project-wide pyproject.toml is applied to all flows through @pypi_base with no additional configuration, which is handy if you want every flow to use the exact same set of dependencies. It is also applied to deployments, unless a deployment declares its own dependencies in config.yml. When the project is deployed, the platform uses Fast Bakery to bake the requirements into a container image automatically.

Assets

Projects track data and models as assets: references built on Metaflow artifacts that add metadata and tracking on top, giving you a model registry and data lineage for the project. By default, obproject-deploy looks for model assets in models/ and data assets in data/. For more information, see Project assets.
If the default asset directory names conflict with existing directories in your project, such as a models/ directory used for data model schemas, customize them in obproject.toml:

Local development

ProjectFlow automatically applies @pypi_base when your project has a pyproject.toml with dependencies. This ensures reproducible environments for both local and remote runs, but requires specifying an environment.

Running flows locally

When @pypi_base is applied, you need to specify an environment:

Skipping dependency isolation

In some contexts, you might want to continue subclassing an obproject.ProjectFlow but turn off the automatic application of @pypi_base. For local iteration using your existing Python environment, you can skip the @pypi_base decorator:
Set OBPROJECT_SKIP_PYPI_BASE per run:
Skipping @pypi_base is convenient when iterating locally. For production deployments through CI/CD, always apply dependencies with --environment=fast-bakery to ensure reproducible builds regardless of your local settings.

CI/CD integration

Projects integrate seamlessly with CI/CD platforms to enable continuous deployment. The obproject-deploy CLI utility available via pip install obproject-utils automates deployment of flows and applications, making it straightforward to set up GitOps workflows.
Starting with ob-project-utils==0.2.35, every flow deployed by obproject-deploy carries a commit-hash:<SHA> tag and a CI-provider-specific run ID tag, such as obproject-deploy-gh-action-run:<ID> for GitHub Actions. Use these tags to trace a running workflow back to the commit and CI build that deployed it. For more information, see Deployment lineage tags.
GitHub Actions can deploy your project automatically when code is pushed to specific branches. Create .github/workflows/deploy.yml:
This workflow:
  • Triggers on pushes to main, develop, and feature branches
  • Authenticates as a machine user with GitHub’s OIDC token
  • Deploys flows, apps, and assets as configured
Use per-component obproject_deploy.toml files to control which branches deploy each app or flow. For details, see Project lifecycle.
If you do not use obproject-deploy, you need to determine when to invoke outerbounds service-principal-configure in your CI runs.

Multi-project repositories

For monorepo setups with multiple independent projects, use obproject_multi.toml at the repository root:
Each project directory contains its own obproject.toml and standard project structure. When you run obproject-deploy from the repository root, it:
  1. Detects obproject_multi.toml
  2. Authenticates as the specified machine user
  3. Deploys each project independently to the configured platform
Individual projects can still be deployed independently by running obproject-deploy from their directories, which will use that project’s specific obproject.toml configuration. Repository structure example:
company-ml-platform
obproject_multi.toml
ml
fraud-detection
obproject.toml
src
models.py
feature_encoders.py
flows
deployments
recommendation
obproject.toml
src
recommenders.py
flows
deployments
pipelines
ingestion
obproject.toml
flows
Each sub-project can have its own src/ directory for shared code. Imports like from models import MyModel work because obproject-deploy sets up PYTHONPATH to include src/ for both flows and deployments. See ob-multi-project-empty for a complete example.

Branch configurations

You can map code branches to different perimeters and deployment configuration files to automate environment isolation. This is configured in obproject.toml:
When you run obproject-deploy, it:
  1. Detects the current git branch
  2. Maps it to an environment using glob pattern matching (first match wins)
  3. Switches to the environment’s perimeter
  4. Uses the environment’s deployment config for applications
This enables workflows like the following:
Environment-specific configurations can vary resources, replicas, and settings:
Branch patterns are matched in the order they appear in [branch_to_environment]. Place specific patterns before wildcards to ensure correct matching.

Flow configs

Flows often use Metaflow’s Config to load JSON configuration files. There are two patterns for organizing configs:
Place config files directly in the flow directory:
flows
train
flow.py
config.json
This works out of the box; no additional configuration needed.
Only register configs in flow_configs when using project-root paths (like configs/model.json). Flow-local configs (like config.json in the same directory as flow.py) do not need registration.
For a complete example with multi-environment deployment and API clients, see the example repository:
outerbounds/ob-project-branch-config
Loading repository data...
To see how these building blocks fit together in a real-world project, continue to Example project.