> ## Documentation Index
> Fetch the complete documentation index at: https://anaconda.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Load .parquet data from S3 to Arrow table

When a Parquet dataset lives in S3, Metaflow's high-throughput S3 client can pull one file or many files into a task at once, and PyArrow can decode them directly into an Arrow table. This page shows the pattern, using a public dataset stored in S3 from [Ookla Global's AWS Open Data Submission](https://registry.opendata.aws/speedtest-global-performance).

<Steps>
  <Step title="Access parquet data in S3">
    Anaconda recommends using `metaflow.S3` in a context manager. The client saves temporary files for the duration of the context, which is why the following example rewrites the file names for access after the scope closes.

    To access one file, use `metaflow.S3.get`. Parquet datasets often have many files, which is a good use case for `s3.get_many`.
  </Step>

  <Step title="Run the flow">
    This flow shows how to:

    * Download multiple Parquet files using `s3.get_many`.
    * Read the result of the first dataset chunk as a [PyArrow](https://arrow.apache.org/docs/python/index.html) table.

    ```py title="load_parquet_to_arrow.py" highlight {21-24} expandable theme={null}
    from metaflow import FlowSpec, step, S3

    BASE_URL = 's3://ookla-open-data/' + \
               'parquet/performance/type=fixed/'
    YEARS = ['2019', '2020', '2021', '2022']
    S3_PATHS = [
        f'year={y}/quarter=1/{y}-' + \
         '01-01_performance_fixed_tiles.parquet' 
        for y in YEARS
    ]

    class ParquetArrowFlow(FlowSpec):
        
        @step
        def start(self):
            self.next(self.load_parquet)

        @step
        def load_parquet(self):
            import pyarrow.parquet as pq
            with S3(s3root=BASE_URL) as s3:
                tmp_data_path = s3.get_many(S3_PATHS)
                first_path = tmp_data_path[0].path
                self.table = pq.read_table(first_path)
            self.next(self.end)

        @step
        def end(self):
            print('Table for first year ' + \
                  f'has shape {self.table.shape}.')

    if __name__ == '__main__':
        ParquetArrowFlow()
    ```

    ```bash theme={null}
    python load_parquet_to_arrow.py run
    ```

    ```text theme={null}
        ...
         [637/end/3308 (pid 7081)] Task is starting.
         [637/end/3308 (pid 7081)] Table for first year has shape (4877036, 7).
         [637/end/3308 (pid 7081)] Task finished successfully.
        ...
    ```
  </Step>

  <Step title="Access artifacts outside of the flow">
    Run the following in any script or notebook to access the contents of the table that was stored as a flow artifact with `self.table`. You can also run quick tests to assert the artifacts have expected properties:

    ```python theme={null}
    from metaflow import Flow
    run = Flow('ParquetArrowFlow').latest_run
    table = run.data.table
    assert run.successful
    assert table.shape == (4877036, 7)
    table.select([1,2,3,4,5])
    ```

    ```text theme={null}
        pyarrow.Table
        tile: string
        avg_d_kbps: int64
        avg_u_kbps: int64
        avg_lat_ms: int64
        tests: int64
    ```
  </Step>
</Steps>
