- Metaflow takes care of packaging your code automatically. You do not have to worry about packaging the project’s Python files.
- Anaconda Platform makes it easier, faster, and more secure to work with 3rd party dependencies and Docker images. For details, see managing dependencies.
Defining an environment for a PyTorch benchmark
This example walks through defining and running a PyTorch benchmark to show how Anaconda Platform handles dependencies. You will run the same benchmark three ways: locally, on cloud compute, and on a GPU.These instructions assume you have installed the
outerbounds package and connected to your platform instance. If you haven’t, start with Connect to the platform and run your first flow.TorchTestFlow, squares a large tensor repeatedly and measures throughput. PyTorch is a good test case for dependency management because it is a large framework with extra requirements beyond the package itself, particularly when GPUs are involved.
Running the benchmark with a declared environment
-
Create a directory for the example and navigate to it:
-
Create a file named
torchtest.pyin that directory with the following contents:The@anaconda_basedecorator declares the environment for every step in the flow. In this case, it specifies Python 3.12 withpytorch. You do not need to build this environment by hand. The platform builds it for you when you run the flow. -
Run the flow. The same file works locally and in the cloud; the only difference is the command.
- Run locally
- Run in the cloud
When you run this command, Metaflow builds the declared environment on your machine and runs the flow inside it.The first run takes a few minutes while packages resolve. The environment is cached for subsequent runs, so the flow starts immediately.For information on using off-the-shelf or custom-built images instead of declaring dependencies, see Managing dependencies.
Running the benchmark with a GPU
Tensor operations are significantly faster on a GPU. To see the difference, run the same benchmark on GPU hardware. GPU tasks need the CUDA driver stack, which belongs in the container image rather than the declared environment. The libraries are large, and they must match the NVIDIA driver on the GPU node. For GPU workloads, use a prebuilt image that includespytorch and CUDA, such as the AWS deep learning containers.
Do I have GPUs in my cluster?Check the availability of GPU instances in your cluster as follows:
- Select Compute in the left-hand navigation.
- Click Pools.
-
Find a compute pool that displays
Has Access to GPUs.
torchtest-gpu.py:
torchtest.py in two ways. First, it adds two decorators to the start step, highlighted above. @resources(gpu=1) requests a GPU, and @gpu_profile() shows a real-time GPU utilization card while the task runs. Second, it omits @anaconda_base, so the task uses the image’s packages instead of building a declared environment.
Run it with the prebuilt image:
- A GPU instance starts to execute the task. You can see auto-scaling events on the Pools tab of the Compute page.
- In the task view, you can observe GPU utilization from
@gpu_profile. - You can also monitor the benchmark results in real time from the custom
@card.