> ## Documentation Index
> Fetch the complete documentation index at: https://anaconda.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Deployments deep dive

export const DefinitionDescription = ({children}) => <dd className="definition-description">{children}</dd>;

export const DefinitionTerm = ({children}) => <dt className="definition-term">{children}</dt>;

export const DefinitionList = ({children}) => <dl className="definition-list">{children}</dl>;

export const GCell = ({children, className}) => <div className={`grid-table-cell ${className || ""}`} role="cell">
    {children}
  </div>;

export const GTH = ({children, className}) => <div className={`grid-table-th ${className || ""}`} role="columnheader">
    {children}
  </div>;

export const GRow = ({children}) => <div className="grid-table-row" role="row">{children}</div>;

export const GBody = ({children}) => <div className="grid-table-body" role="rowgroup">{children}</div>;

export const GHead = ({children}) => <div className="grid-table-head" role="rowgroup">{children}</div>;

export const GTable = ({children, className, cols}) => <div className={`grid-table not-prose overflow-hidden rounded-2xl ${className || ""}`} style={{
  "--grid-table-cols": cols
}} role="table">
    {children}
  </div>;

The [Writing your first deployment](/docs/platform/guides/inference/writing-your-first-deployment) tutorial gets a service up and running. This page goes deeper into the features available when working with deployments, along with best practices and tips.

## Commands quick start

The CLI exposes commands for working with deployments:

<GTable cols="35% 65%">
  <GHead>
    <GRow>
      <GTH>Command</GTH>
      <GTH>Description</GTH>
    </GRow>
  </GHead>

  <GBody>
    <GRow>
      <GCell>[`outerbounds app deploy`](/docs/platform/cli/app/deploy)</GCell>
      <GCell>Create a new deployment, or modify an existing deployment with the same `name`.</GCell>
    </GRow>

    <GRow>
      <GCell>[`outerbounds app list`](/docs/platform/cli/app/list)</GCell>
      <GCell>List all the deployments on the platform.</GCell>
    </GRow>

    <GRow>
      <GCell>[`outerbounds app info`](/docs/platform/cli/app/info)</GCell>
      <GCell>Show details for a deployment.</GCell>
    </GRow>

    <GRow>
      <GCell>[`outerbounds app logs`](/docs/platform/cli/app/logs)</GCell>
      <GCell>Fetch logs from a deployment's workers.</GCell>
    </GRow>

    <GRow>
      <GCell>[`outerbounds app delete`](/docs/platform/cli/app/delete)</GCell>
      <GCell>Delete a deployment by its `--name`.</GCell>
    </GRow>
  </GBody>
</GTable>

## Setting up the deployment config

While you *can* pass all your configuration options as CLI flags, Anaconda recommends using config files for the following reasons:

<DefinitionList>
  <DefinitionTerm>Visibility</DefinitionTerm>
  <DefinitionDescription>All your deployment settings in one easily accessible place.</DefinitionDescription>
  <DefinitionTerm>Version control</DefinitionTerm>
  <DefinitionDescription>Committing the version control to the repository naturally tracks configuration changes over time.</DefinitionDescription>
  <DefinitionTerm>Reusability</DefinitionTerm>
  <DefinitionDescription>Replicate deployments across environments.</DefinitionDescription>
  <DefinitionTerm>Less error-prone</DefinitionTerm>
  <DefinitionDescription>Avoid typos in long CLI commands.</DefinitionDescription>
  <DefinitionTerm>Explainability</DefinitionTerm>
  <DefinitionDescription>Annotate fields with long multi-line comments in a config file more easily than in a shell command.</DefinitionDescription>
</DefinitionList>

<Tip>
  For rapid prototyping, you can override the settings defined in your config file with CLI command `[OPTIONS]`. See the [Deployments CLI reference](/docs/platform/cli/app/deploy) for the available options.
</Tip>

The following config file contains every field covered on this page. The top block is required; the rest are commented out. Uncomment a block and fill in the values to include it in your settings.

```yaml title="config.yaml" expandable theme={null}
# Required. The name is the unique identifier for the deployment.
# Port and commands are required unless you deploy a proxy capsule or use
# a pre-built image with no_deps (see Dependency management).
name: my-service
port: 8000
commands:
- python app.py

# Authentication: Browser, API, or BrowserAndAPI
# auth:
#   type: API

# Environment variables available to the app
# environment:
#   DOWNLOAD_DIR: /tmp/models
#   MODEL_NAME: llm_name

# Secrets from platform integrations, exposed as environment variables
# secrets:
# - openai-api-key

# Extra file suffixes to include in the code package
# package:
#   suffixes:
#   - .sql
#   - .txt

# Resource reservations
# resources:
#   cpu: "2"
#   memory: "8Gi"
#   gpu: "1"
#   disk: "100Gi"
#   shared_memory: "2Gi"

# Scaling: either a fixed count or autoscaling bounds
# replicas:
#   min: 1
#   max: 10
#   scaling_policy:
#     rpm: 100

# Compute pool targeting
# compute_pools:
# - gpu-pool-1
# - gpu-pool-2

# Dependencies: requirements file, package lists, or a custom image
# dependencies:
#   python: "3.11"
#   from_requirements_file: requirements.txt

# URL behavior
# generate_static_url: true
# url_slug: my-custom-url

# Storage: connect the deployment to the platform PostgreSQL database
# persistence: postgres
```

## Configuring environment variables

If your deployment depends on environment variables, you can use the `environment:` top-level field in the config file to define them:

```yaml theme={null}
environment:
  DOWNLOAD_DIR: /tmp/models
  MODEL_NAME: llm_name
```

## Configuring secrets

For sensitive information that your deployment depends on, such as API keys, Anaconda recommends using integrations. Integrations let you store secrets and use them safely in your Metaflow tasks and deployments. For instructions on creating an integration as a secret, see [Configuring secrets](/docs/platform/guides/security/configuring-secrets).

Once you have configured a secret, reference it in your deployment by name:

```yaml theme={null}
secrets: 
- openai-api-key # Must match the name of an integration
```

The config file contains only the integration name, never a credential value. When the deployment's containers start, the platform fetches the integration's key-value pairs and injects each key as an environment variable. To rotate a credential, update the integration; a restarted pod picks up the new value without a rebuild or redeploy.

Any keys defined inside the integration are then available as environment variables on your deployment. For example, if you set up a key called `MY_API_KEY` inside an integration you named `openai-api-key`, use it in your deployment with:

```python theme={null}
api_key = os.environ.get("MY_API_KEY")
```

## Packaging non-Python files

By default, the platform packages all of the Python files on your local system so they can run on the cloud, replicating the folder structure so relative paths stay the same.

Just like Metaflow tasks, you can define an additional list of file suffixes to include in your deployment:

```yaml theme={null}
package:
    suffixes:
    - .sql
    - .txt
```

## Multi-step startups

Sometimes you have bootstrap scripts that must run before your deployment starts. A good example is a `model-downloader.py` that downloads a model to a specified location, followed by an `app.py` that loads the downloaded model and powers inference on it.

Set this up with the `commands` field:

```yaml theme={null}
commands:
  - "python model_downloader.py --model_name $MODEL_NAME"
  - "vllm serve $DOWNLOAD_DIR/$MODEL_NAME --dtype=half --task score"
```

## Access types for deployments

You might want to serve UI apps such as Streamlit or TensorBoard, or API endpoints such as Flask, FastAPI, or vLLM apps.

If you set up UI access, anyone who has access to the platform UI has access to your deployment. The deployment is guarded by the same auth that guards the platform UI.

If you set up API access, the endpoint is accessible over API for programmatic clients. Callers access the endpoint by providing their Metaflow token as `x-api-key`.

Control this setting with the `auth` field:

```yaml theme={null}
auth:
    type: Browser # UI access. Use 'API' for API access.
```

Sometimes a deployment must do both: serve UI routes on some paths and API endpoints on others. For example, Flask and FastAPI apps often expose a route map on one path and API endpoints on all others. In that case, use:

```yaml theme={null}
auth:
    type: BrowserAndAPI
```

With `BrowserAndAPI`, the platform mints two separate URLs:

1. UI routes are served at `ui-c-123.<YOUR_DEPLOYMENT_DOMAIN>`.
2. API routes are served at `api-c-123.<YOUR_DEPLOYMENT_DOMAIN>`.

## Resource management

To make sure your deployments perform as expected, configure the right resources. Reserve resources for your deployment with the `resources` field:

```yaml theme={null}
resources:
  cpu: "2"              # CPU cores
  memory: "8Gi"         # Memory (use Mi or Gi units)
  gpu: "1"              # Number of GPUs
  disk: "100Gi"         # Persistent storage
  shared_memory: "2Gi"  # Shared memory (useful for vLLM, Ray, etc.)
```

## Scaling workers

Different deployments have different usage patterns and requirements. Some deployments have predictable traffic, while others have variable traffic. Your requirements might also change depending on whether your deployment is for testing and prototyping or for production.

You can run a fixed number of workers, or set up autoscaling based on request volume.

### Using a fixed number of workers

Set a fixed number of workers that never autoscale:

```yaml theme={null}
replicas:
    fixed: 3
```

This keeps three workers available at all times. If a worker encounters an error, the platform automatically replaces it to maintain your configured worker count.

Fixed worker counts work well when:

* You have steady-state traffic that does not vary much.
* Your SLAs are strict and cannot afford delays when responding to requests.

### Autoscaling workers

You can also set up autoscaling based on the request rate per minute:

```yaml theme={null}
replicas:
  min: 1              # Minimum workers
  max: 10             # Maximum workers
  scaling_policy:
    rpm: 100          # Scale up at 100 requests/minute per worker
```

In this example, if you see \~500 requests per minute, the platform runs five workers automatically. The workers scale down when the request rate drops. Setting `min: 0` enables scaling to zero when the deployment is idle.

Autoscaling works well when:

* You have variable or unpredictable traffic patterns.
* You want to control costs.
* You do not have strict latency requirements. Autoscaling workers takes some time, depending on the type of compute instances they use.

## Targeting compute pools

You might want your deployment to run on specific compute pools. Reasons include:

* **Cost tracking**: With a separate compute pool carved out for your deployment, you can calculate the cost of running the deployment from the cost incurred on that pool.
* **Compute isolation**: Isolate critical applications from other deployments, workstations, and tasks so other workloads cannot impact them.
* **Compute requirements**: When using GPUs, not all instances are the same. You might want to target a particular class of GPU.

Pin your deployment to one or more compute pools with the `compute_pools` field:

```yaml theme={null}
compute_pools:
- gpu-pool-1
- gpu-pool-2
```

<Warning>
  A compute pool can only run deployments if the setting is enabled on the pool. On the **Compute** page, select a compute pool (or create a new one) and select "Inference deployments" under Advanced Routing.
</Warning>

## Authenticating for cloud access

Your deployment might have dependencies in your cloud account that it needs to operate, such as an app that reads from your S3 buckets or DynamoDB tables to serve requests.

A deployment automatically runs with the default task role of its [perimeter](/docs/platform/concepts/what-is-a-perimeter), so by default it has access to everything a Metaflow task running in that perimeter can access.

To override the default role, set the `OBP_AWS_DEPLOYMENT_IDENTITY` environment variable in your config to the role you want to use. The role must be properly tagged and assumable by the task role, as described in [Configuring secrets](/docs/platform/guides/security/configuring-secrets#using-a-custom-iam-role).

## Dependency management

By default, if you have a `requirements.txt` at the root of your project, the platform uses it to bake a Docker image with all the specified packages.

You can also explicitly point your deployment at a particular requirements file:

```yaml theme={null}
dependencies: 
    python: "3.11" # Python version to use in your built Docker container
    from_requirements_file: requirements.txt
  # from_pyproject_toml: pyproject.toml 
```

Just like Metaflow tasks, you can also define your dependencies by specifying PyPI or conda packages directly:

```yaml theme={null}
dependencies: 
    python: "3.11" 
    pypi: 
        numpy: 1.23.0
        pandas: ''
```

```yaml theme={null}
dependencies: 
    python: "3.11" 
    conda: 
        numpy: 1.23.0
        pandas: ''
```

To provide your own Docker image instead:

```yaml theme={null}
image: python:3.10-slim
```

You can optionally declare that you want to use the image as provided, without installing any packages on top of it:

```yaml theme={null}
image: python:3.10-slim
no_deps: true
```

## Managing URLs

By default, every deployment gets a random identifier as part of its URL. If you create, destroy, and recreate an identical app, the recreated app gets a different URL.

To keep the URL consistent across redeployments with the same project, branch, and name in the same perimeter, set `generate_static_url: true` in your config.

### Defining your own subdomain

You cannot define a fully custom subdomain for your deployment, but you can define a custom slug that appears in the subdomain, accompanied by the `ui-` or `api-` prefix.

Set `url_slug: my-custom-url` in your config. If no other deployment has taken the slug, your deployment is minted the URL `ui-my-custom-url.<YOUR_DEPLOYMENT_DOMAIN>` or `api-my-custom-url.<YOUR_DEPLOYMENT_DOMAIN>`, depending on the auth type.

## Connecting to your PostgreSQL database

The platform provisions a PostgreSQL database inside your cloud account. While it mostly serves as the home for recording metadata about your Metaflow runs, you can also use it as the storage layer for your deployment.

This is particularly useful for use cases like hyperparameter optimization with Optuna, which uses a relational database to record experiment metadata.

To connect your deployment to the platform's PostgreSQL database, set:

```yaml theme={null}
persistence: postgres
```

You can then connect to the database at `localhost:5432` inside your deployment, using your `METAFLOW_SERVICE_AUTH_KEY` as the database password. The platform injects `METAFLOW_SERVICE_AUTH_KEY` into your deployment's environment automatically, so no setup is required to read it.

## Monitoring

Open the **Deployments** page in the platform UI and select your deployment to find:

* Logs for all workers, useful for debugging or general sanity checks.
* Metrics for all workers, to understand resource tuning.
* Metrics for the entire deployment (request rates, latencies), to understand performance.
* Autoscaling charts, to see how your deployment is scaling.
* The general health of your deployment and its workers.
* Configuration attributes and the update history of your deployment.
