> ## Documentation Index
> Fetch the complete documentation index at: https://anaconda.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Load CSV data in Metaflow steps

When your flow needs data from a CSV file, you can bundle the file with the flow using `IncludeFile`. The file's contents become available to every step, locally and on remote compute. This page shows how to read a CSV into a flow, transform the data, and store the result as a flow artifact you can access later from any script or notebook.

<Steps>
  <Step title="Acquire the CSV">
    This example uses a CSV from the [Metaflow tutorials](https://docs.metaflow.org/getting-started/tutorials/season-1-the-local-experience/episode00), downloaded by the `save_data_locally` function defined outside of the flow.
  </Step>

  <Step title="Run the flow">
    This flow shows how to:

    * Include a CSV saved locally for all steps in the flow.
    * Add a feature to each data point.
    * Save the new data as a flow artifact.

    ```py title="load_csv_data.py" expandable theme={null}
    from metaflow import FlowSpec, step, IncludeFile
    import pandas as pd

    def save_data_locally():
        url = "https://raw.githubusercontent.com/" + \
              "Netflix/metaflow/master/metaflow"
        data_path = "/tutorials/02-statistics/movies.csv"
        local_path = "./movies.csv"
        df = pd.read_csv(url+data_path)
        df.to_csv(local_path)

    class CSVFlow(FlowSpec):
        
        data = IncludeFile("data", default="./movies.csv")
        
        @step
        def start(self):
            self.next(self.use_csv)
            
        @step
        def use_csv(self):
            import pandas as pd 
            from io import StringIO
            df = pd.read_csv(StringIO(self.data),
                             index_col=0)
            f = lambda x: x < 2000
            df["is_before_2000"] = df["title_year"].apply(f)
            self.new_df = df
            self.next(self.end)
            
        @step
        def end(self):
            result = self.new_df.is_before_2000.sum() 
            print(f"Number of pre-2000 movies is {result}.")
            
    if __name__ == "__main__":
        save_data_locally()
        CSVFlow()
    ```

    ```bash theme={null}
    python load_csv_data.py run
    ```

    ```text theme={null}
        ...
         [1654221300950244/end/3 (pid 71595)] Task is starting.
         [1654221300950244/end/3 (pid 71595)] Number of pre-2000 movies is 1023.
         [1654221300950244/end/3 (pid 71595)] Task finished successfully.
        ...
    ```
  </Step>

  <Step title="Access artifacts outside of the flow">
    Run the following in any script or notebook to access the contents of the dataframe that was stored as a flow artifact with `self.new_df`:

    ```python theme={null}
    from metaflow import Flow 
    run = Flow("CSVFlow").latest_run
    assert run.successful
    run.data.new_df.is_before_2000.sum()
    ```

    ```text theme={null}
        1023
    ```
  </Step>
</Steps>
