> ## Documentation Index
> Fetch the complete documentation index at: https://anaconda.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Load .parquet data from S3 to pandas dataframe

When a Parquet dataset lives in S3, Metaflow's high-throughput S3 client can pull one file or many files into a task at once, and pandas can read the result straight into a dataframe for analysis. This page shows the pattern, using a public dataset stored in S3 from [Ookla Global's AWS Open Data Submission](https://registry.opendata.aws/speedtest-global-performance).

<Steps>
  <Step title="Access parquet data in S3">
    To access one file, use `metaflow.S3.get`. Parquet datasets often have many files, which is a good use case for `s3.get_many`.
  </Step>

  <Step title="Run the flow">
    This flow shows how to:

    * Download multiple Parquet files using `s3.get_many`.
    * Read the result of one year of the dataset as a pandas dataframe.

    ```py title="load_parquet_to_pandas.py" highlight {21-24} expandable theme={null}
    from metaflow import FlowSpec, step, S3

    BASE_URL = 's3://ookla-open-data/' + \
                'parquet/performance/type=fixed/'
    YEARS = ['2019', '2020', '2021', '2022']
    S3_PATHS = [
        f'year={y}/quarter=1/{y}-' + \
         '01-01_performance_fixed_tiles.parquet' 
        for y in YEARS
    ]

    class ParquetPandasFlow(FlowSpec):

        @step
        def start(self):
            self.next(self.load_parquet)

        @step
        def load_parquet(self):
            import pandas as pd
            with S3(s3root=BASE_URL) as s3:
                tmp_data_path = s3.get_many(S3_PATHS)
                first_path = tmp_data_path[0].path
                self.df = pd.read_parquet(first_path)
            self.next(self.end)

        @step
        def end(self):
            print('DataFrame for first year ' + \
                  f'has shape {self.df.shape}.')

    if __name__ == '__main__':
        ParquetPandasFlow()
    ```

    ```bash theme={null}
    python load_parquet_to_pandas.py run
    ```

    ```text theme={null}
        ...
         [638/end/3312 (pid 7120)] Task is starting.
         [638/end/3312 (pid 7120)] DataFrame for first year has shape (4877036, 7).
         [638/end/3312 (pid 7120)] Task finished successfully.
        ...
    ```
  </Step>

  <Step title="Access artifacts outside of the flow">
    Run the following in any script or notebook to access the contents of the dataframe that was stored as a flow artifact with `self.df`:

    ```python theme={null}
    from metaflow import Flow
    Flow('ParquetPandasFlow').latest_run.data.df.head()
    ```

    ```text theme={null}
                quadkey                                              tile  avg_d_kbps  avg_u_kbps  avg_lat_ms  tests  devices
    0  0231113112003202  POLYGON((-90.6591796875 38.4922941923613, -90....       66216       12490          13     28        4
    1  1322111021111001  POLYGON((110.352172851562 21.2893743558604, 11...      102598       37356          13     15        4
    2  3112203030003110  POLYGON((138.592529296875 -34.9219710361638, 1...       24686       18736          18    162      106
    3  0320000130321312  POLYGON((-87.637939453125 40.225024210605, -87...       17674       13989          78    364        4
    4  0320001332313103  POLYGON((-84.7430419921875 38.9209554204673, -...      441192      218955          22     14        1
    ```
  </Step>
</Steps>
