> ## Documentation Index
> Fetch the complete documentation index at: https://anaconda.com/docs/llms.txt
> Use this file to discover all available pages before exploring further.

# Compare LLM approaches for text classification

This tutorial walks you through building a text classification system that infers whether product reviewers are likely to recommend a product. You will compare zero-shot LLM inference (OpenAI) to a hybrid approach (embeddings plus XGBoost) on streaming data from a Postgres database.

<Note>
  Creating a resource integration requires an administrator role. If you do not have administrator access, ask your administrator to create the integration before you begin.
</Note>

By the end of this tutorial, you will have:

* An OpenAI API key configured as a platform integration
* A workstation notebook that validates your setup
* A baseline flow that scores review sentiment using zero-shot LLM inference
* A sensor flow that monitors the database for new data
* Baseline and candidate deployments that process reviews automatically
* A monitoring dashboard that compares workflow performance over time

## Configure the OpenAI API key

Several steps in this tutorial call the OpenAI API. You need an API key to run the code samples.

To configure the key as a platform integration:

1. Select **Integrations** in the left-hand navigation.
2. Click **OpenAI** in the **Add an Integration** section.
3. Enter a name for the integration.
4. Enter a description for the integration.
5. Enter your OpenAI API key.
6. Click **Add**.

## Download the tutorial content

Download the tutorial content to your workstation:

```bash theme={null}
outerbounds tutorials pull --url https://outerbounds-journeys-content.s3.us-west-2.amazonaws.com/main/journeys.tar.gz --destination-dir ~/learn
```

<Tip>
  This command downloads all tutorial content as a single bundle. If you've already worked through other tutorials, you likely already have this and do not need to run the command again.
</Tip>

The LLM tutorial content is in `~/learn/llm-end-to-end`. If you prefer a different location, replace `~/learn` with a directory of your choice.

## Validate your setup

Open the notebook in `00-setup-nb` from the `~/learn/llm-end-to-end` directory. This notebook installs the required dependencies and validates your OpenAI API key.

## Explore the dataset

Open the notebook in `01-baseline-nb` from the `~/learn/llm-end-to-end` directory. This notebook introduces the dataset and calls OpenAI to score the sentiment of clothing reviews. If you do not have an OpenAI API key, you can read along without running the cells.

## Run the baseline evaluation

Evaluate how well the LLM performs zero-shot inference on historical data with known labels. Run this command:

```bash theme={null}
cd ~/learn/llm-end-to-end/02-baseline-flow
python flow.py --environment=fast-bakery run --with kubernetes --eval True --n 200
```

## Run inference on new reviews

Use the same flow to score new reviews without labels:

```bash theme={null}
cd ~/learn/llm-end-to-end/02-baseline-flow
python flow.py --environment=fast-bakery run --with kubernetes
```

## Understand the data pipeline

In production ML systems, there is often a delay between inference time and when the true label becomes available. In this tutorial, the marketing team might wait days or weeks before knowing whether a customer actually recommended the product.

Key details about the data pipeline:

* New reviews appear in the database at regular intervals (every 10 minutes in this simulation).
* The `recommended_ind` column uses NULL values for reviews without feedback yet. For existing labels, this column is `1` if the reviewer recommended the product, and `0` if they did not.
* To work with historical data that has labels, use `fetch_table(table_name, only_labeled=True)`.
* To work with the most recent batch of unlabeled data, use `fetch_table(table_name, only_labeled=False)`.

This pattern lets you train and evaluate on historical data with known outcomes while making predictions on new data where feedback is pending.

## Deploy the sensor flow

Deploy a sensor flow that monitors the database and sends an event to the platform when new data is available:

```bash theme={null}
cd ~/learn/llm-end-to-end/03-sensor-flow
python flow.py --environment=fast-bakery argo-workflows-create
```

You can adjust the monitoring interval by changing the `schedule` parameter in the `argo-workflows-create` command.

## Deploy the baseline flow

Deploy the baseline flow so it listens to the sensor flow's event and runs inference when new data is available:

```bash theme={null}
cd ~/learn/llm-end-to-end/04-deploy-baseline
python flow.py --production --environment=fast-bakery argo-workflows-create
```

## Deploy a candidate flow

Deploy a candidate flow that uses a different LLM model to compare against the baseline:

```bash theme={null}
cd ~/learn/llm-end-to-end/05-deploy-candidate
python flow.py --production --environment=fast-bakery argo-workflows-create
```

## Monitor and compare flows

Open the notebook in `06-monitoring` from the `~/learn/llm-end-to-end` directory. This notebook uses the Metaflow Client API to compare the performance of the candidate and baseline flows across branches and over time.

## Next steps

To build on this tutorial:

* Explore other LLM providers and model sizes.
* Integrate the classification pipeline into your existing ML workflows.
