Skip to main content
When a pandas dataframe grows too large to process comfortably in one piece, you can split it into chunks and write each chunk to its own Parquet file. This page shows a pattern for doing that in Metaflow: Apache Arrow’s zero-copy slicing to create the chunks without copying data, combined with Metaflow’s foreach to process the chunks in parallel branches.
To read a dataframe from S3 into memory, see Load .parquet data from S3 to pandas dataframe.
1

Gather data

Suppose you have curated a dataset:
Your goal is to store this data efficiently in Parquet files.
2

Determine how to chunk the data

This example uses PyArrow to split the dataframe into chunks, as shown in this utility function that the flow uses:
dataframe_utils.py
3

Run the flow

This flow shows how to load the data into a pandas dataframe and apply the following steps:
  • Use the pyarrow.Table.from_pandas method to load the data to Arrow memory.
  • In parallel branches:
    • Use pyarrow.Table.slice to make zero-copy views of chunks of the table.
    • Apply a transformation to the table. In this case, it appends a column.
    • Move the chunks to your S3 bucket using pyarrow.parquet.write_table.
  • Pick a chunk and verify the existence of the new transformed column.
chunk_dataframe.py