# Deploy MLflow with the mlflow template

Source: https://docs.quake.ai/resources/deployments/deploy-mlflow-template
Markdown: https://docs.quake.ai/resources/deployments/deploy-mlflow-template.md

---

# Deploy MLflow with the mlflow template

Stand up [MLflow](https://mlflow.org), an open-source experiment-tracking and model-registry platform, on a single Quake AI instance using the [validated OpenTofu template](/docs/platform/validation#how-infrastructure-templates-are-checked) `mlflow`. You apply the template, wire it to an external PostgreSQL database and an Object Storage bucket, start the tracking server, and log a test run from your workstation.

MLflow is the tracking layer for ML platform teams. You run it yourself; this is a self-hosted tool you operate, not a managed service.

<Figure size="md" caption="What you'll build: an MLflow tracking server on a single instance, with metadata in PostgreSQL and artifacts in Object Storage">

```d2
direction: right

dev: You {shape: person}
domain: Your domain\n(DNS A record)
fip: Floating IP
postgres: PostgreSQL\nmetadata
bucket: Object Storage\nartifacts
instance: Ubuntu instance {
  caddy: Caddy\nreverse proxy
  mlflow: MLflow\ntracking server
  caddy -> mlflow: proxies 443 to 5000
}

dev -> domain: HTTPS UI
dev -> mlflow: log runs\nport 5000
domain -> fip
fip -> instance.caddy
mlflow -> postgres: metadata
mlflow -> bucket: artifacts
```

</Figure>

<PricingCompanion
  components={[
    { kind: "template", slug: "mlflow", required: true },
  ]}
/>

## Prerequisites

You need:

- OpenTofu 1.6.0 or later (or Terraform 1.6.0 or later) installed locally.
- Your OpenStack credentials sourced into the shell (`source openrc.sh`). See [the OpenStack CLI guide](/docs/tools/openstack-cli).
- An SSH keypair that already exists in your project. Record its name for the `key_name` variable.
- A copy of the `mlflow` template directory from [the template reference page](/resources/iac-templates/mlflow).
- An external PostgreSQL database with a database and user for MLflow metadata. The [self-managed PostgreSQL template](/resources/iac-templates/self-managed-postgres) is one path; record the private IP for `postgres_host`.
- An Object Storage bucket and S3 credentials. The [S3 Storage with ACLs template](/resources/iac-templates/s3-storage-acl) provisions both; record the bucket name, endpoint URL, access key, and secret key.
- Your workstation's public IP address, so you can open the UI port to it for first-boot setup. Find it with `curl -sS https://api.ipify.org`.

A domain is optional for first boot. You add it in step 4 to serve the UI over HTTPS.

## Step 1: Set the variables and apply the template

The tracking UI listens on port 5000 over plain HTTP. The template's security group restricts port 5000 to `ui_allowed_cidr`, which defaults to the private network only, so the raw UI stays off the public internet. To reach the UI from your workstation for first-boot setup, set `ui_allowed_cidr` to your own address.

Copy the template's example variables file and open it:

```bash
cp terraform.tfvars.example terraform.tfvars
```

Set the required values:

```hcl
key_name          = "YOUR_KEY_NAME"
postgres_host     = "POSTGRES_PRIVATE_IP"
artifact_bucket   = "YOUR_BUCKET_NAME"
s3_endpoint       = "https://us-east-1.rumble.cloud"
ui_allowed_cidr   = "YOUR_IP/32"
```



If you would rather not expose port 5000 at all, leave `ui_allowed_cidr` at its default and reach the UI over an SSH tunnel instead: `ssh -L 5000:localhost:5000 ubuntu@YOUR_FLOATING_IP`, then open `http://localhost:5000`. Once you add a domain in step 4, Caddy serves the UI over HTTPS on port 443 and you no longer need port 5000 open.



Initialize the working directory, preview the plan, and apply:

```bash
tofu init
tofu plan
tofu apply
```

OpenTofu provisions a private network, a router, a security group, a block volume mounted at `/var/lib/docker`, an instance, and a floating IP. On first boot, cloud-init mounts the data volume and installs Docker Engine. MLflow does not start until you add credentials in step 2.

When the apply finishes, read the outputs:

```bash
tofu output
```

Record `floating_ip` and `tracking_url`.

## Step 2: Add credentials and start the tracking server

No credential ships with this template. SSH to the instance and edit `/opt/mlflow/.env`. Uncomment and set the password and S3 credentials:

```bash
ssh ubuntu@YOUR_FLOATING_IP
sudo nano /opt/mlflow/.env
```

Add values for `POSTGRES_PASSWORD`, `AWS_ACCESS_KEY_ID`, and `AWS_SECRET_ACCESS_KEY`. The PostgreSQL user must already have access to the database named in `POSTGRES_DB`. Create the database on your Postgres instance first if it does not exist:

```sql
CREATE DATABASE mlflow;
CREATE USER mlflow WITH PASSWORD 'your-secure-password';
GRANT ALL PRIVILEGES ON DATABASE mlflow TO mlflow;
```

Build the MLflow image and start the server:

```bash
cd /opt/mlflow
sudo docker compose up -d --build
sudo docker compose ps
```

Open `tracking_url` (for example `http://YOUR_FLOATING_IP:5000`) in your browser. The MLflow home page loads when the container is healthy.



Until you attach a domain in step 4, the UI is served over unencrypted HTTP on port 5000, reachable only from `ui_allowed_cidr`. Avoid sending production credentials over it from a shared or public network. Adding a domain (step 4) moves the UI to HTTPS on port 443.



## Step 3: Log a test experiment

From your workstation, install the MLflow client and point it at the tracking server:

```bash
pip install mlflow
export MLFLOW_TRACKING_URI=http://YOUR_FLOATING_IP:5000
```

Log a short run:

```python
import mlflow

mlflow.set_experiment("quake-smoke-test")

with mlflow.start_run(run_name="hello-quake"):
    mlflow.log_param("framework", "cpu-only")
    mlflow.log_metric("accuracy", 0.91)
    mlflow.set_tag("source", "deployment-walkthrough")
```

Refresh the MLflow UI in your browser. The `quake-smoke-test` experiment appears with the run you logged.

This walkthrough logs metadata only. To log a model artifact, call `mlflow.sklearn.log_model` or `mlflow.log_artifact` in the same run; MLflow writes the file to the Object Storage bucket you configured.

## Step 4: Serve the UI over HTTPS with Caddy

The template leaves ports 80 and 443 open for a reverse proxy. [Caddy](https://caddyserver.com) obtains and renews a TLS certificate automatically once a domain resolves to the instance.

1. Create a DNS **A record** for your domain (for example `mlflow.example.com`) pointing at `YOUR_FLOATING_IP`. Follow [How to point a domain at a Quake AI resource](/docs/network/how-to/point-domain-to-quake-ai). Wait until the record resolves:

```bash
dig +short mlflow.example.com
```

2. SSH to the instance and create `/opt/mlflow/Caddyfile`:

```text
mlflow.example.com {
  reverse_proxy 127.0.0.1:5000
}
```

3. Add Caddy to `/opt/mlflow/docker-compose.yml`:

```yaml
services:
  caddy:
    image: caddy:2
    restart: unless-stopped
    network_mode: host
    volumes:
      - /opt/mlflow/Caddyfile:/etc/caddy/Caddyfile
      - caddy_data:/data
volumes:
  caddy_data:
```

4. Apply the change:

```bash
cd /opt/mlflow
sudo docker compose up -d
```

Open `https://mlflow.example.com` and confirm the padlock. Once HTTPS works, close direct access to port 5000 by setting `ui_allowed_cidr` back to the private network in `terraform.tfvars` and running `tofu apply`.

Update training jobs and notebooks to use `MLFLOW_TRACKING_URI=https://mlflow.example.com`.

## What you built

- **Applied the `mlflow` template** to provision a network, security group, data volume, instance, and floating IP
- **Wired PostgreSQL and Object Storage** by setting connection variables in tfvars and credentials in `/opt/mlflow/.env`
- **Started the MLflow tracking server** and logged a test experiment from your workstation
- **Served the UI over HTTPS** by pointing a domain at the floating IP and routing it through a Caddy reverse proxy

## Scope of this deployment

This template runs a single-VM MLflow tracking host, not a managed experiment-tracking cloud. The instance is CPU-only and runs in one region. You operate the instance, Docker, MLflow, the external PostgreSQL database, and the Object Storage bucket yourself: back them up, patch them, and watch storage growth as run volume increases.

This template hosts experiment tracking and the model registry only. GPU training and fine-tuning run on an external backend you operate. Point those jobs at your GPU cluster and set `MLFLOW_TRACKING_URI` to this server so metrics and artifacts still land in your registry.

## Next steps

- [MLflow template reference](/resources/iac-templates/mlflow): parameters, metadata and artifact wiring, and resource map
- [self-managed PostgreSQL template](/resources/iac-templates/self-managed-postgres): the metadata store this template points at
- [S3 Storage with ACLs template](/resources/iac-templates/s3-storage-acl): the artifact bucket this template points at
- [Deploy JupyterHub with the jupyterhub template](/resources/deployments/deploy-jupyterhub-template): multi-user notebooks that log runs to MLflow
- [How to store application secrets and inject them at runtime](/docs/security/how-to/inject-app-secrets): move database and S3 credentials out of plain environment files

## Clean up

When you no longer need the deployment, destroy everything the template created:

```bash
tofu destroy
```

The external PostgreSQL database and Object Storage bucket are not destroyed by this command; delete those resources separately if you no longer need them. Remove the DNS A record you created in step 4.
