Skip to main content
The Writing your first deployment tutorial gets a service up and running. This page goes deeper into the features available when working with deployments, along with best practices and tips.

Commands quick start

The CLI exposes commands for working with deployments:

Setting up the deployment config

While you can pass all your configuration options as CLI flags, Anaconda recommends using config files for the following reasons:
For rapid prototyping, you can override the settings defined in your config file with CLI command [OPTIONS]. See the Deployments CLI reference for the available options.
The following config file contains every field covered on this page. The top block is required; the rest are commented out. Uncomment a block and fill in the values to include it in your settings.
config.yaml

Configuring environment variables

If your deployment depends on environment variables, you can use the environment: top-level field in the config file to define them:

Configuring secrets

For sensitive information that your deployment depends on, such as API keys, Anaconda recommends using integrations. Integrations let you store secrets and use them safely in your Metaflow tasks and deployments. For instructions on creating an integration as a secret, see Configuring secrets. Once you have configured a secret, reference it in your deployment by name:
The config file contains only the integration name, never a credential value. When the deployment’s containers start, the platform fetches the integration’s key-value pairs and injects each key as an environment variable. To rotate a credential, update the integration; a restarted pod picks up the new value without a rebuild or redeploy. Any keys defined inside the integration are then available as environment variables on your deployment. For example, if you set up a key called MY_API_KEY inside an integration you named openai-api-key, use it in your deployment with:

Packaging non-Python files

By default, the platform packages all of the Python files on your local system so they can run on the cloud, replicating the folder structure so relative paths stay the same. Just like Metaflow tasks, you can define an additional list of file suffixes to include in your deployment:

Multi-step startups

Sometimes you have bootstrap scripts that must run before your deployment starts. A good example is a model-downloader.py that downloads a model to a specified location, followed by an app.py that loads the downloaded model and powers inference on it. Set this up with the commands field:

Access types for deployments

You might want to serve UI apps such as Streamlit or TensorBoard, or API endpoints such as Flask, FastAPI, or vLLM apps. If you set up UI access, anyone who has access to the platform UI has access to your deployment. The deployment is guarded by the same auth that guards the platform UI. If you set up API access, the endpoint is accessible over API for programmatic clients. Callers access the endpoint by providing their Metaflow token as x-api-key. Control this setting with the auth field:
Sometimes a deployment must do both: serve UI routes on some paths and API endpoints on others. For example, Flask and FastAPI apps often expose a route map on one path and API endpoints on all others. In that case, use:
With BrowserAndAPI, the platform mints two separate URLs:
  1. UI routes are served at ui-c-123.<YOUR_DEPLOYMENT_DOMAIN>.
  2. API routes are served at api-c-123.<YOUR_DEPLOYMENT_DOMAIN>.

Resource management

To make sure your deployments perform as expected, configure the right resources. Reserve resources for your deployment with the resources field:

Scaling workers

Different deployments have different usage patterns and requirements. Some deployments have predictable traffic, while others have variable traffic. Your requirements might also change depending on whether your deployment is for testing and prototyping or for production. You can run a fixed number of workers, or set up autoscaling based on request volume.

Using a fixed number of workers

Set a fixed number of workers that never autoscale:
This keeps three workers available at all times. If a worker encounters an error, the platform automatically replaces it to maintain your configured worker count. Fixed worker counts work well when:
  • You have steady-state traffic that does not vary much.
  • Your SLAs are strict and cannot afford delays when responding to requests.

Autoscaling workers

You can also set up autoscaling based on the request rate per minute:
In this example, if you see ~500 requests per minute, the platform runs five workers automatically. The workers scale down when the request rate drops. Setting min: 0 enables scaling to zero when the deployment is idle. Autoscaling works well when:
  • You have variable or unpredictable traffic patterns.
  • You want to control costs.
  • You do not have strict latency requirements. Autoscaling workers takes some time, depending on the type of compute instances they use.

Targeting compute pools

You might want your deployment to run on specific compute pools. Reasons include:
  • Cost tracking: With a separate compute pool carved out for your deployment, you can calculate the cost of running the deployment from the cost incurred on that pool.
  • Compute isolation: Isolate critical applications from other deployments, workstations, and tasks so other workloads cannot impact them.
  • Compute requirements: When using GPUs, not all instances are the same. You might want to target a particular class of GPU.
Pin your deployment to one or more compute pools with the compute_pools field:
A compute pool can only run deployments if the setting is enabled on the pool. On the Compute page, select a compute pool (or create a new one) and select “Inference deployments” under Advanced Routing.

Authenticating for cloud access

Your deployment might have dependencies in your cloud account that it needs to operate, such as an app that reads from your S3 buckets or DynamoDB tables to serve requests. A deployment automatically runs with the default task role of its perimeter, so by default it has access to everything a Metaflow task running in that perimeter can access. To override the default role, set the OBP_AWS_DEPLOYMENT_IDENTITY environment variable in your config to the role you want to use. The role must be properly tagged and assumable by the task role, as described in Configuring secrets.

Dependency management

By default, if you have a requirements.txt at the root of your project, the platform uses it to bake a Docker image with all the specified packages. You can also explicitly point your deployment at a particular requirements file:
Just like Metaflow tasks, you can also define your dependencies by specifying PyPI or conda packages directly:
To provide your own Docker image instead:
You can optionally declare that you want to use the image as provided, without installing any packages on top of it:

Managing URLs

By default, every deployment gets a random identifier as part of its URL. If you create, destroy, and recreate an identical app, the recreated app gets a different URL. To keep the URL consistent across redeployments with the same project, branch, and name in the same perimeter, set generate_static_url: true in your config.

Defining your own subdomain

You cannot define a fully custom subdomain for your deployment, but you can define a custom slug that appears in the subdomain, accompanied by the ui- or api- prefix. Set url_slug: my-custom-url in your config. If no other deployment has taken the slug, your deployment is minted the URL ui-my-custom-url.<YOUR_DEPLOYMENT_DOMAIN> or api-my-custom-url.<YOUR_DEPLOYMENT_DOMAIN>, depending on the auth type.

Connecting to your PostgreSQL database

The platform provisions a PostgreSQL database inside your cloud account. While it mostly serves as the home for recording metadata about your Metaflow runs, you can also use it as the storage layer for your deployment. This is particularly useful for use cases like hyperparameter optimization with Optuna, which uses a relational database to record experiment metadata. To connect your deployment to the platform’s PostgreSQL database, set:
You can then connect to the database at localhost:5432 inside your deployment, using your METAFLOW_SERVICE_AUTH_KEY as the database password. The platform injects METAFLOW_SERVICE_AUTH_KEY into your deployment’s environment automatically, so no setup is required to read it.

Monitoring

Open the Deployments page in the platform UI and select your deployment to find:
  • Logs for all workers, useful for debugging or general sanity checks.
  • Metrics for all workers, to understand resource tuning.
  • Metrics for the entire deployment (request rates, latencies), to understand performance.
  • Autoscaling charts, to see how your deployment is scaling.
  • The general health of your deployment and its workers.
  • Configuration attributes and the update history of your deployment.