
Flows
Flows refer to Metaflow flows, often interconnected through events. They form the backbone of your projects, handling data processing and ETL, model training and finetuning, autonomous and batch inference, and any other types of background processing and high-performance computing. In projects, flows are stored under a subdirectoryflows, one Metaflow flow (named flow.py) per subdirectory, alongside any supporting Python modules and packages. As a best practice, it is useful to add a README.md file for each flow describing its role. They will be surfaced in the UI as well.
Writing a ProjectFlow
Importantly, project flows should subclass from ProjectFlow instead of Metaflow’s standard FlowSpec. In other words, simply author your flows like this:
BaseFlow pattern to enrich flows with functionality related to the project structure. Besides this small detail, you can use all Metaflow features in your flows.
A typical flow hierarchy in a project repository ends up looking like this:
flows
etl
flow.py
README.md
feature_transformations.py
sql
process_data.sql
train_model
flow.py
README.md
model.py
Deployments
Deployments are microservices that serve requests through real-time APIs. Common use cases:Flows and deployments working together
The platform’s strength comes from the tight connection between flows and deployments, bridging the offline and online worlds:- A flow can continuously update a database for RAG, which a deployed agent then uses in real time.
- A deployed app can monitor model performance and trigger a retraining flow.
- A flow can deploy a model endpoint programmatically, for instance, whenever a new model has been trained.
Deployments in a project
In your project, place deployments in thedeployments directory. Each deployment is defined by a configuration file, config.yml, as documented in Deployments deep dive. You can define dependencies for the deployment in a standard requirements.txt or pyproject.toml. As with flows, it is recommended to add a README.md for each deployment.
The project hierarchy looks like this:
deployments
monitoring_dashboard
streamlit_app.py
config.yml
pyproject.toml
README.md
model_endpoint
fastapi_server.py
config.yml
pyproject.toml
README.md
support_agent
agent.py
config.yml
pyproject.toml
README.md
Deployment commands
Whenobproject-deploy deploys apps, it runs from the project root directory, so commands in your config.yml must use paths relative to the project root, not the deployment directory.
For example, if your project structure is:
my_project
deployments
dashboard
app.py
config.yml
api
main.py
config.yml
src
shared_utils.py
config.yml commands with paths relative to the project root:
- Shared imports: Apps can import modules from
src/directly (from my_module import ...). - Consistency: All paths are relative to the same root, making them predictable.
Code
Effective management of software dependencies is essential for building production-quality projects and enabling rapid iteration and collaboration. A project has one platform configuration file,obproject.toml, and one or more dependency manifests, such as pyproject.toml or requirements.txt. The platform configuration says what the project is and how it deploys; the dependency manifests say what packages the code needs:
A typical project consists of multiple layers of software dependencies:
- Code defining flows and deployments, organized into subdirectories.
- Project-level shared libraries under the
srcdirectory. - Organization-level libraries shared across projects.
- Third-party dependencies, such as
pandasandtorch, declared at the flow, deployment, or project level.
fraud_detection_model
obproject.toml
pyproject.toml
README.md
src
feature_encoders
__init__.py
feature_encoder.py
flows
trainer
flow.py
mymodel.py
README.md
deployments
inference
fastapi_server.py
config.yml
README.md
Code for flows and deployments
In addition to the entrypoint file (flow.py) or deployment server, each flow or deployment can include supporting modules and packages, such as mymodel.py in the tree above.
Project-level shared libraries
Place libraries shared within a project under thesrc directory, as packages. For example, a feature_encoders package used both during training and inference ensures offline-online consistency of features:
src
feature_encoders
__init__.py
feature_encoder.py
__init__.py, include:
obproject-deploy, it automatically sets up PYTHONPATH so your flows and apps can import these modules directly, such as from feature_encoders import MyEncoder.
Organization-level libraries
Libraries shared across projects can be handled in two ways:- If you can set
METAFLOW_PACKAGE_POLICYin packages, simplypip installthem as usual or add them to yourPYTHONPATH. Once youimportthem in your flows and deployments, they get packaged automatically. This is a convenient option for private packages, even if they are notpip install-able from a package repository. - If the shared libraries are pushed to a package repository, private or public, treat them like third-party dependencies.
Declaring dependencies
You can declare dependencies at three levels in a project:- Per flow or step: Use Metaflow’s
@pypior@condadecorators in your flow code. - Per deployment: Add a
requirements.txtorpyproject.tomlto the deployment directory, or declare dependencies in itsconfig.yml. - Project-wide: Place a
pyproject.tomlat the root of the project next toobproject.toml. For example:
pyproject.toml is applied to all flows through @pypi_base with no additional configuration, which is handy if you want every flow to use the exact same set of dependencies. It is also applied to deployments, unless a deployment declares its own dependencies in config.yml.
When the project is deployed, the platform uses Fast Bakery to bake the requirements into a container image automatically.
Assets
Projects track data and models as assets: references built on Metaflow artifacts that add metadata and tracking on top, giving you a model registry and data lineage for the project. By default,obproject-deploy looks for model assets in models/ and data assets in data/. For more information, see Project assets.
Local development
ProjectFlow automatically applies @pypi_base when your project has a pyproject.toml with dependencies. This ensures reproducible environments for both local and remote runs, but requires specifying an environment.
Running flows locally
When@pypi_base is applied, you need to specify an environment:
Skipping dependency isolation
In some contexts, you might want to continue subclassing anobproject.ProjectFlow but turn off the automatic application of @pypi_base. For local iteration using your existing Python environment, you can skip the @pypi_base decorator:
- Environment variable
- Shell profile
- Project config
Set
OBPROJECT_SKIP_PYPI_BASE per run:CI/CD integration
Projects integrate seamlessly with CI/CD platforms to enable continuous deployment. Theobproject-deploy CLI utility available via pip install obproject-utils automates deployment of flows and applications, making it straightforward to set up GitOps workflows.
Starting with
ob-project-utils==0.2.35, every flow deployed by obproject-deploy carries a commit-hash:<SHA> tag and a CI-provider-specific run ID tag, such as obproject-deploy-gh-action-run:<ID> for GitHub Actions. Use these tags to trace a running workflow back to the commit and CI build that deployed it. For more information, see Deployment lineage tags.- GitHub Actions
- Azure DevOps
- CircleCI
- GitLab CI/CD
GitHub Actions can deploy your project automatically when code is pushed to specific branches. Create This workflow:
.github/workflows/deploy.yml:- Triggers on pushes to main, develop, and feature branches
- Authenticates as a machine user with GitHub’s OIDC token
- Deploys flows, apps, and assets as configured
If you do not use
obproject-deploy, you need to determine when to invoke outerbounds service-principal-configure in your CI runs.Multi-project repositories
For monorepo setups with multiple independent projects, useobproject_multi.toml at the repository root:
obproject.toml and standard project structure. When you run obproject-deploy from the repository root, it:
- Detects
obproject_multi.toml - Authenticates as the specified machine user
- Deploys each project independently to the configured platform
obproject-deploy from their directories, which will use that project’s specific obproject.toml configuration.
Repository structure example:
company-ml-platform
obproject_multi.toml
ml
fraud-detection
obproject.toml
src
models.py
feature_encoders.py
flows
deployments
recommendation
obproject.toml
src
recommenders.py
flows
deployments
pipelines
ingestion
obproject.toml
flows
src/ directory for shared code. Imports like from models import MyModel work because obproject-deploy sets up PYTHONPATH to include src/ for both flows and deployments.
See ob-multi-project-empty for a complete example.
Branch configurations
You can map code branches to different perimeters and deployment configuration files to automate environment isolation. This is configured inobproject.toml:
obproject-deploy, it:
- Detects the current git branch
- Maps it to an environment using glob pattern matching (first match wins)
- Switches to the environment’s perimeter
- Uses the environment’s deployment config for applications
Branch patterns are matched in the order they appear in
[branch_to_environment]. Place specific patterns before wildcards to ensure correct matching.Flow configs
Flows often use Metaflow’sConfig to load JSON configuration files. There are two patterns for organizing configs:
- Flow-local configs
Place config files directly in the flow directory:This works out of the box; no additional configuration needed.
flows
train
flow.py
config.json
outerbounds/ob-project-branch-config
Loading repository data...
To see how these building blocks fit together in a real-world project, continue to Example project.