Skip to main content
A deployment is a long-running service that stays available to answer requests, rather than running to completion like a run. Deployments run on compute in your data plane, governed by the perimeter they belong to.

What are deployments for?

Some work can’t run to completion. A model that answers requests, an API that other applications call, a dashboard people check throughout the day: these need to stay running and reachable. A deployment keeps a service like this available, answering requests as they arrive, until you stop it. Deployments cover more than model inference. Anything that needs to stay available to answer requests can run as a deployment, whether it’s a web app, an agent endpoint, or a container serving a large language model.

How deployments work

A deployment runs as one or more identical copies of your service, called workers. Multiple workers let the deployment answer more requests at once and stay available if one fails. You choose how many workers a deployment has, and deployments can autoscale to handle variable traffic, including scaling to zero when idle. Like flows and workstations, deployment workers run on the compute pools assigned to their perimeter, so a perimeter must have compute available for them. Because a deployment is governed by its perimeter, the same access rules and policies that apply to runs and workstations apply to it as well. You manage deployments and monitor their status, logs, and metrics through the platform.

Working with deployments

For the full guide on creating and managing deployments, see Deploy a model from the model catalog.