Skip to main content
Now that we have seen the basic building blocks of executing code in the cloud, the next step is to understand how to manage and run more complex projects. In particular, we need to package software dependencies, both our own code and any 3rd party libraries, for reliable cloud execution. Dependency management in Python is known to be a headache (in particular affecting ML/AI projects that have extra requirements due to large packages and GPU drivers), but Anaconda Platform streamlines the process while allowing you to choose an approach that works for your needs. Specifically:
  • Metaflow takes care of packaging your code automatically. You do not have to worry about packaging the project’s Python files.
  • Anaconda Platform makes it easier, faster, and more secure to work with 3rd party dependencies and Docker images. For details, see managing dependencies.

Defining an environment for a PyTorch benchmark

This example walks through defining and running a PyTorch benchmark to show how Anaconda Platform handles dependencies. You will run the same benchmark three ways: locally, on cloud compute, and on a GPU.
These instructions assume you have installed the outerbounds package and connected to your platform instance. If you haven’t, start with Connect to the platform and run your first flow.
The benchmark flow, TorchTestFlow, squares a large tensor repeatedly and measures throughput. PyTorch is a good test case for dependency management because it is a large framework with extra requirements beyond the package itself, particularly when GPUs are involved.

Running the benchmark with a declared environment

  1. Create a directory for the example and navigate to it:
  2. Create a file named torchtest.py in that directory with the following contents:
    The @anaconda_base decorator declares the environment for every step in the flow. In this case, it specifies Python 3.12 with pytorch. You do not need to build this environment by hand. The platform builds it for you when you run the flow.
  3. Run the flow. The same file works locally and in the cloud; the only difference is the command.
    When you run this command, Metaflow builds the declared environment on your machine and runs the flow inside it.The first run takes a few minutes while packages resolve. The environment is cached for subsequent runs, so the flow starts immediately.
    For information on using off-the-shelf or custom-built images instead of declaring dependencies, see Managing dependencies.

Running the benchmark with a GPU

Tensor operations are significantly faster on a GPU. To see the difference, run the same benchmark on GPU hardware. GPU tasks need the CUDA driver stack, which belongs in the container image rather than the declared environment. The libraries are large, and they must match the NVIDIA driver on the GPU node. For GPU workloads, use a prebuilt image that includes pytorch and CUDA, such as the AWS deep learning containers.
Do I have GPUs in my cluster?Check the availability of GPU instances in your cluster as follows:
  1. Select Compute in the left-hand navigation.
  2. Click Pools.
  3. Find a compute pool that displays Has Access to GPUs.
    The Pools tab of the Compute page. Has access to a GPU is spotlighted.
If you don’t have access to a compute pool with GPUs, contact your administrator.
The GPU run needs its own version of the flow. Save this as torchtest-gpu.py:
This version differs from torchtest.py in two ways. First, it adds two decorators to the start step, highlighted above. @resources(gpu=1) requests a GPU, and @gpu_profile() shows a real-time GPU utilization card while the task runs. Second, it omits @anaconda_base, so the task uses the image’s packages instead of building a declared environment. Run it with the prebuilt image:
The first GPU run takes a few minutes while the cluster launches a new GPU instance and pulls the image. When the flow runs:
  1. A GPU instance starts to execute the task. You can see auto-scaling events on the Pools tab of the Compute page.
  2. In the task view, you can observe GPU utilization from @gpu_profile.
  3. You can also monitor the benchmark results in real time from the custom @card.
In this example, the GPU handles 106 squarings per second. A CPU-only laptop manages about 2. Next, learn how to take a flow from development to production in Deploying to production.