foreach to process the chunks in parallel branches.
To read a dataframe from S3 into memory, see Load .parquet data from S3 to pandas dataframe.
1
Gather data
Suppose you have curated a dataset:Your goal is to store this data efficiently in Parquet files.
2
Determine how to chunk the data
This example uses PyArrow to split the dataframe into chunks, as shown in this utility function that the flow uses:
dataframe_utils.py
3
Run the flow
This flow shows how to load the data into a pandas dataframe and apply the following steps:
- Use the
pyarrow.Table.from_pandasmethod to load the data to Arrow memory. - In parallel branches:
- Use
pyarrow.Table.sliceto make zero-copy views of chunks of the table. - Apply a transformation to the table. In this case, it appends a column.
- Move the chunks to your S3 bucket using
pyarrow.parquet.write_table.
- Use
- Pick a chunk and verify the existence of the new transformed column.
chunk_dataframe.py