Skip to main content
When a Parquet dataset lives in S3, Metaflow’s high-throughput S3 client can pull one file or many files into a task at once, and PyArrow can decode them directly into an Arrow table. This page shows the pattern, using a public dataset stored in S3 from Ookla Global’s AWS Open Data Submission.
1

Access parquet data in S3

Anaconda recommends using metaflow.S3 in a context manager. The client saves temporary files for the duration of the context, which is why the following example rewrites the file names for access after the scope closes.To access one file, use metaflow.S3.get. Parquet datasets often have many files, which is a good use case for s3.get_many.
2

Run the flow

This flow shows how to:
  • Download multiple Parquet files using s3.get_many.
  • Read the result of the first dataset chunk as a PyArrow table.
load_parquet_to_arrow.py
3

Access artifacts outside of the flow

Run the following in any script or notebook to access the contents of the table that was stored as a flow artifact with self.table. You can also run quick tests to assert the artifacts have expected properties: