> ## Documentation Index
> Fetch the complete documentation index at: https://anaconda.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Deploy a model from the model catalog

export const Comments = ({children}) => {
  return <div class="my-4 px-5 py-4 overflow-hidden rounded-2xl flex gap-3 border border-zinc-500/20 bg-zinc-50/50 dark:border-zinc-500/30 dark:bg-zinc-500/10" data-callout-type="comments">
      <div class="w-4">
        <svg width="14" height="14" viewBox="0 0 640 640" fill="currentColor" xmlns="http://www.w3.org/2000/svg" class="w-5 h-5" aria-label="Comments">
            <path d="M320 112C434.9 112 528 205.1 528 320C528 434.9 434.9 528 320 528C205.1 528 112 434.9 112 320C112 205.1 205.1 112 320 112zM320 576C461.4 576 576 461.4 576 320C576 178.6 461.4 64 320 64C178.6 64 64 178.6 64 320C64 461.4 178.6 576 320 576zM280 400C266.7 400 256 410.7 256 424C256 437.3 266.7 448 280 448L360 448C373.3 448 384 437.3 384 424C384 410.7 373.3 400 360 400L352 400L352 312C352 298.7 341.3 288 328 288L280 288C266.7 288 256 298.7 256 312C256 325.3 266.7 336 280 336L304 336L304 400L280 400zM320 256C337.7 256 352 241.7 352 224C352 206.3 337.7 192 320 192C302.3 192 288 206.3 288 224C288 241.7 302.3 256 320 256z" />
        </svg>
      </div>
      <div class="text-sm prose min-w-0 w-full">
        {children}
      </div>
    </div>;
};

The fastest way to serve a model on Anaconda Platform is to deploy it directly from the model catalog. Deployments from the catalog serve the model as-is, with no application logic around it, which makes them well suited for evaluating and comparing models before you commit to one.

The platform packages the model into a long-running service, schedules it on an inference compute pool, and gives you an authenticated API endpoint and a built-in chat interface to try it out, with no code or configuration files required.

To build a service around a model instead, see [Writing your first deployment](/docs/platform/guides/inference/writing-your-first-deployment). For background on how deployments work, see [What is a deployment?](/docs/platform/concepts/what-is-deployment)

For more inference patterns, including CLI-based deployment for vLLM (safetensors) models and the `@vllm` and `@llamacpp` decorators, see the inference examples repository:

<GitHub.Repo repo="outerbounds/inference-examples" />

## Deploying a model

1. In the left-hand navigation, select **Resources**, then select **Models**.

2. Confirm the context picker shows the perimeter you want to deploy into. The model catalog shows the models available in your current perimeter, as determined by that perimeter's model policies.

   <Frame>
     <img src="https://mintcdn.com/anaconda-29683c67/cwLLt5dtVRvj7OPF/images/platform/plat_deploy_model_perimeter.png?fit=max&auto=format&n=cwLLt5dtVRvj7OPF&q=85&s=9ce0b58af88950fcaf0583ed4e28c4cd" alt="The perimeter context picker open on the model details page, showing the default perimeter selected" width="1866" height="663" data-path="images/platform/plat_deploy_model_perimeter.png" />
   </Frame>

3. Select a model from the list to open its details page.

4. Click **Deploy Model**.

   <Frame>
     <img src="https://mintcdn.com/anaconda-29683c67/cwLLt5dtVRvj7OPF/images/platform/plat_deploy_model_button.png?fit=max&auto=format&n=cwLLt5dtVRvj7OPF&q=85&s=888c2a13bba39bf9779a3768628bb2a2" alt="The model details page with the Deploy Model button" width="1866" height="666" data-path="images/platform/plat_deploy_model_button.png" />
   </Frame>

5. In the Deploy model dialog, confirm the **Project** context for the deployment.

6. Under **Resources**, set the CPU, memory, and disk for the deployment. Use the model's **file size** and **estimated RAM** values shown under the selected file as a guide. Memory must exceed the estimated RAM, and disk must exceed the file size with room for the serving runtime.

7. Select a compute pool for the deployment, or leave the selection empty to let the platform schedule it on the best-fitting pool. If no compute pool can manage your requested resources, choose a smaller model file or reduce the values for the resources you are requesting.

   <Frame>
     <img src="https://mintcdn.com/anaconda-29683c67/VD0yQ0tXYWIdTsBU/images/platform/plat_deploy_model_details.png?fit=max&auto=format&n=VD0yQ0tXYWIdTsBU&q=85&s=d34c9db3bdcb88a2e3be9c7de40855b1" alt="The Deploy model dialog showing Project, Model, and File fields with resource inputs for CPU, memory, disk, GPU, and shared memory, and a list of compute pools" width="1866" height="1082" data-path="images/platform/plat_deploy_model_details.png" />
   </Frame>

8. Click **Create**.

The platform opens the **Deployments** view. The deployment registers within a few seconds, then takes several minutes to provision and start.

<Note>
  If the Deployments view shows "Deployment not found" immediately after creating a deployment, refresh the page. The deployment list updates as the platform registers the new deployment.
</Note>

## Accessing the deployment

When the deployment is ready, it appears in the Deployments list with a green status indicator. Select it to open the details view, which shows:

* The **API endpoint URL**, labeled "available at". This is the endpoint your applications call.
* **Additional UI routes**. For certain model file types (`.gguf` models served with `llamacpp`), this includes a built-in chat interface for trying the model interactively. Safetensors models served with vLLM expose the API endpoint only. Both the endpoint and the chat UI require visitors to be signed in to the platform.
* The serving image, compute pool, and resources the deployment is running with.
* Charts for requests per minute and autoscaling activity.

<Frame>
  <img src="https://mintcdn.com/anaconda-29683c67/VD0yQ0tXYWIdTsBU/images/platform/plat_deployments_model_details.png?fit=max&auto=format&n=VD0yQ0tXYWIdTsBU&q=85&s=bb7fc7404f0e8f7738029bc1668adb84" alt="The deployment details view for a deployed model showing the endpoint URL, additional UI routes, serving image, compute pool, and requests per minute chart" width="1866" height="1082" data-path="images/platform/plat_deployments_model_details.png" />
</Frame>

The deployment name gets a random suffix (for example, `gemma-2-2b-it-dm2kkh`) so you can deploy the same model more than once without name conflicts.

## Trying the model

The deployment's API endpoint is the primary interface: send requests to the endpoint URL with your application or tools like `curl`. For gguf models served with llamacpp, the deployment also includes a built-in chat interface.

To open it, click the **Additional UI routes** link in the deployment details. You will be asked to sign in if you are not already signed in. From the chat UI, you can send messages to the model and inspect its responses without writing any code. Each response shows its token count, latency, and throughput in tokens per second, which gives you a quick read on how the deployment performs under your sizing choices.

<Note>
  The chat interface is available only for `.gguf` models served with `llamacpp`. Safetensors models served with `vLLM` expose the API endpoint only.
</Note>

<Frame>
  <img src="https://mintcdn.com/anaconda-29683c67/VD0yQ0tXYWIdTsBU/images/platform/plat_deployed_model_chat.png?fit=max&auto=format&n=VD0yQ0tXYWIdTsBU&q=85&s=2d7e743dd87c910c6ead85ee3ba57b6f" alt="The chat interface for a deployed model showing a message input field with the model name" width="1866" height="577" data-path="images/platform/plat_deployed_model_chat.png" />
</Frame>

## Deleting the deployment

To tear down a deployment you no longer need, use the CLI:

```sh theme={null}
outerbounds app delete --name <DEPLOYMENT_NAME>
```

<Comments>
  Replace \<DEPLOYMENT\_NAME> with the full name of the deployment as shown in the Deployments list, including the random suffix (for example, `gemma-2-2b-it-dm2kkh`).
</Comments>

Deletion is asynchronous. The deployment might show a terminating state briefly before it disappears from the list.

## Next steps

* To deploy a custom service (such as a FastAPI app or a fine-tuned model) instead of a catalog model, see [Writing your first deployment](/docs/platform/guides/inference/writing-your-first-deployment).
* For deployment configuration options in depth, see [Deployments deep dive](/docs/platform/guides/inference/deployments-deep-dive).
* To manage deployments from the command line, see the [Deployments CLI reference](/docs/platform/cli/app/deploy).
