How the pieces fit together
Cost optimization on the platform comes down to how a handful of concepts relate:- Each step of a workflow executes as a task. A task’s size is determined by its resource requirements (memory, CPU, GPU, and disk), defined with the
@resourcesdecorator. Workstations are a special kind of task that consume resources in the same way. - The platform launches cloud instances to execute tasks and terminates them automatically as demand changes, packing tasks onto instances to keep them busy.
- Each instance belongs to a compute pool, a group of instances of certain types.
- Compute pools can be shared across perimeters, the isolated environments that separate teams or environments such as production and development.
Optimizing cost
A key observation about cloud costs is that you pay for every second an instance is running, not for every second it is doing useful work. The main lever for optimizing costs is therefore increasing instance utilization, which minimizes wasted instance-seconds.Right-sizing resource requests
Imagine a task that loads a dataframe into memory. You estimate the dataframe needs at most 100GB of RAM, so you annotate the task with@resources(memory=100_000). To execute the task, the platform needs an instance with at least 100GB of memory, such as an r5.4xlarge instance with 128GB of RAM.
In practice, the task might occupy only 70GB of RAM, leading to a situation like this:
Optimizing instance types
Another optimization opportunity in this scenario is using smaller instance types. The platform bin-packs tasks onto instances efficiently, so a larger instance can execute multiple smaller tasks simultaneously. Smaller instance types therefore do not always lead to higher utilization, since more instances might be required to handle the same load. Usually, the best approach is to start with instances large enough to handle all workloads. The system collects instrumentation about the total utilization rate over time, which you can observe in the Nodes report. If certain instance types later prove suboptimal, you can swap them for other types.Leveraging spot instances, reservations, discounts, and multiple clouds
Anaconda Platform works with any instance types available in your cloud account. You can use spot instances, instance reservations, negotiated discounts, and credits to further lower compute costs. These resources are typically configured as a dedicated compute pool in your cluster. In addition, the platform makes it easy to bring in compute pools from other clouds, such as resources from GCP when you primarily use AWS. This allows you to leverage credits, discounts, and other incentives across clouds, further lowering the total cost of compute.Configuring compute poolsContact Anaconda support to configure compute pools that use spot instances and reservations, and to learn about available incentives for moving compute between clouds.
Leveraging shared compute pools
Another source of underutilization affects systems with multiple compute pools and perimeters. Imagine a typical setup with two perimeters, a Production environment and a Development environment, each with its own dedicated compute pool,prod-aws-main and dev-aws-main respectively: