> ## Documentation Index
> Fetch the complete documentation index at: https://anaconda.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Overview of cost optimization

Cloud compute is a relatively inexpensive commodity, at least compared to the cost of human experts. For that reason, converting compute cycles into human productivity is generally a good investment. For example, an ML developer who needs to train multiple models can train them in parallel rather than sequentially, trading a modest increase in compute cost for hours of saved time.

The same logic applies in production. When you deploy a new version of a workflow, you can run it alongside the existing production version instead of replacing it immediately. Running both in parallel lets you verify the new version's results before you cut over to it. This doubles compute requirements temporarily, but the cost of incorrect results reaching production would be orders of magnitude higher.

Anaconda Platform is designed to give you cost-efficient access to compute, so you can choose approaches that optimize for business outcomes without worrying about runaway infrastructure costs. For more background, see [The 6 Steps to Cost-Optimized ML/AI/Data Workloads](https://www.anaconda.com/blog/six-steps-to-cost-optimization).

## How the pieces fit together

Cost optimization on the platform comes down to how a handful of concepts relate:

* Each step of a [workflow](/docs/platform/concepts/what-is-a-workflow) executes as a [task](/docs/platform/concepts/what-is-a-workload). A task's size is determined by its resource requirements (memory, CPU, GPU, and disk), defined with [the `@resources` decorator](https://docs.metaflow.org/scaling/remote-tasks/requesting-resources). [Workstations](/docs/platform/concepts/what-is-a-workstation) are a special kind of task that consume resources in the same way.
* The platform launches cloud instances to execute tasks and terminates them automatically as demand changes, packing tasks onto instances to keep them busy.
* Each instance belongs to a [compute pool](/docs/platform/concepts/what-is-compute), a group of instances of certain types.
* Compute pools can be shared across [perimeters](/docs/platform/concepts/what-is-a-perimeter), the isolated environments that separate teams or environments such as production and development.

<Frame>
  <img src="https://mintcdn.com/anaconda-29683c67/TXgc_5FN8574WLab/images/platform/migrated/cost-concepts.svg?fit=max&auto=format&n=TXgc_5FN8574WLab&q=85&s=4c5942f1b52d4a17a245aa3e8b1e6f78" alt="Diagram showing Metaflow tasks packed onto cloud instances, with instances grouped into compute pools and pools shared across perimeters" width="1920" height="1080" data-path="images/platform/migrated/cost-concepts.svg" />
</Frame>

In the diagram, green boxes denote Metaflow tasks and purple boxes denote cloud instances.

## Optimizing cost

<Tip>
  **Beware of premature cost optimization.**

  Before worrying about cost, check the **Historical spend** report to see your actual costs. You might find the total is low enough that it does not warrant further optimization.
</Tip>

A key observation about cloud costs is that **you pay for every second an instance is running, not for every second it is doing useful work**. The main lever for optimizing costs is therefore increasing instance utilization, which minimizes wasted instance-seconds.

### Right-sizing resource requests

Imagine a task that loads a dataframe into memory. You estimate the dataframe needs at most 100GB of RAM, so you annotate the task with `@resources(memory=100_000)`. To execute the task, the platform needs an instance with at least 100GB of memory, such as an `r5.4xlarge` instance with 128GB of RAM.

In practice, the task might occupy only 70GB of RAM, leading to a situation like this:

<Frame>
  <img src="https://mintcdn.com/anaconda-29683c67/TXgc_5FN8574WLab/images/platform/migrated/cost-task-resources.svg?fit=max&auto=format&n=TXgc_5FN8574WLab&q=85&s=f8c09094e717acb154a71b1f9f2a8b5c" alt="Diagram of a single task using 70GB of memory on an instance with 128GB, leaving the remainder underutilized" width="1920" height="1080" data-path="images/platform/migrated/cost-task-resources.svg" />
</Frame>

While the task executes, at least 30GB of RAM is underutilized, and up to 58GB if no other tasks fit on the same instance simultaneously. Worse, the effect is often multiplied across many instances, leading to significant underutilization:

<Frame>
  <img src="https://mintcdn.com/anaconda-29683c67/TXgc_5FN8574WLab/images/platform/migrated/cost-many-tasks.svg?fit=max&auto=format&n=TXgc_5FN8574WLab&q=85&s=fdd45149f78b5c968030a0eddcb5fc32" alt="Diagram of many tasks spread across many instances, each instance partially underutilized" width="1920" height="1080" data-path="images/platform/migrated/cost-many-tasks.svg" />
</Frame>

Resource underutilization is common in distributed computing systems like AWS EMR and Databricks. This inefficiency is difficult to detect and often goes unnoticed, resulting in unnecessarily high compute costs that are not accurately attributed to inefficient tasks.

The platform has a report specifically for this problem. Open the **Flows** report to see how each task has utilized its resources historically. For more information, see [Using cost reports](/docs/platform/guides/monitor/using-cost-reports).

#### Optimizing instance types

Another optimization opportunity in this scenario is using smaller instance types. The platform bin-packs tasks onto instances efficiently, so a larger instance can execute multiple smaller tasks simultaneously. Smaller instance types therefore do not always lead to higher utilization, since more instances might be required to handle the same load.

Usually, the best approach is to start with instances large enough to handle all workloads. The system collects instrumentation about the total utilization rate over time, which you can observe in the **Nodes** report. If certain instance types later prove suboptimal, you can swap them for other types.

<Tip>
  **Don't try to optimize instance types prematurely.**

  Unless you are working with specialized and expensive instance types, such as large GPU instances, run actual workloads on the platform before optimizing your instance mix. Optimizing instance types is more effective when based on real utilization data.
</Tip>

#### Leveraging spot instances, reservations, discounts, and multiple clouds

Anaconda Platform works with any instance types available in your cloud account. You can use spot instances, instance reservations, negotiated discounts, and credits to further lower compute costs. These resources are typically configured as a dedicated compute pool in your cluster.

In addition, the platform makes it easy to [bring in compute pools from other clouds](/docs/platform/guides/compute/running-steps-across-clouds), such as resources from GCP when you primarily use AWS. This allows you to leverage credits, discounts, and other incentives across clouds, further lowering the total cost of compute.

<Note>
  **Configuring compute pools**

  Contact Anaconda support to configure compute pools that use spot instances and reservations, and to learn about available incentives for moving compute between clouds.
</Note>

### Leveraging shared compute pools

Another source of underutilization affects systems with multiple compute pools and perimeters. Imagine a typical setup with two perimeters, a **Production** environment and a **Development** environment, each with its own dedicated compute pool, `prod-aws-main` and `dev-aws-main` respectively:

<Frame>
  <img src="https://mintcdn.com/anaconda-29683c67/TXgc_5FN8574WLab/images/platform/migrated/cost-separate-pools.svg?fit=max&auto=format&n=TXgc_5FN8574WLab&q=85&s=768ddf999fb0062035c6e841589872bf" alt="Diagram of production and development perimeters with separate compute pools, where the production pool is oversubscribed while development instances sit idle" width="1920" height="1080" data-path="images/platform/migrated/cost-separate-pools.svg" />
</Frame>

In this scenario, the production pool faces heavy demand, causing tasks to queue because the pool lacks capacity to run all tasks simultaneously. Meanwhile, the development pool has an idle instance and an underutilized one.

Separate compute pools can be beneficial to ensure development workloads never consume resources needed for production. However, the strict boundary results in suboptimal resource allocation and usage.

Alternatively, one compute pool can be shared between the two perimeters:

<Frame>
  <img src="https://mintcdn.com/anaconda-29683c67/TXgc_5FN8574WLab/images/platform/migrated/cost-unified-pool.svg?fit=max&auto=format&n=TXgc_5FN8574WLab&q=85&s=e368e735e96af36d1421a72e80e777f8" alt="Diagram of production and development perimeters sharing a single compute pool, with instances fully utilized" width="1920" height="1080" data-path="images/platform/migrated/cost-unified-pool.svg" />
</Frame>

In this case, resources are allocated on the fly to the perimeter that needs them most, leading to higher throughput, higher utilization, and lower total cost. Contact Anaconda support to set up perimeters and their compute pools.
