# MLflow experiment tracking

Source: https://docs.quake.ai/resources/iac-templates/mlflow
Markdown: https://docs.quake.ai/resources/iac-templates/mlflow.md

---

# MLflow experiment tracking

This pattern composes Compute, Network, Block Storage, and Object Storage into a self-hosted experiment-tracking and model-registry server you run on infrastructure you control.

## What this template does

Provisions a single instance running [MLflow](https://mlflow.org), an open-source experiment-tracking and model-registry platform (a self-hosted alternative to Weights & Biases or Neptune). You log parameters, metrics, and artifacts from notebooks or training jobs, compare runs in the UI, and promote models through the registry:

- Compute instance that runs the MLflow tracking server in Docker, sized for moderate concurrent logging (2 vCPU and 2 GiB RAM)
- Private network, subnet, router, port, and security group; a floating IP for public access
- A block volume mounted at `/var/lib/docker`, so Docker state lives on a volume you can grow rather than on the boot disk
- cloud-init installs Docker Engine and writes the MLflow compose stack on first boot
- External PostgreSQL for experiment metadata and Object Storage for run artifacts (models, plots, files)

MLflow is the tracking layer for ML platform teams. It records experiments and hosts the model registry on a VM you own, which keeps metadata and artifact paths on your infrastructure.

No credential ships with this template. You add the PostgreSQL password and Object Storage credentials to `/opt/mlflow/.env` on the instance after apply.

## Parameters

| Parameter | Description | Default |
| --- | --- | --- |
| `key_name` | SSH keypair name (must already exist) | No default |
| `postgres_host` | PostgreSQL host for metadata (for example a [self-managed PostgreSQL](/resources/iac-templates/self-managed-postgres) private IP) | No default |
| `artifact_bucket` | Object Storage bucket for run artifacts | No default |
| `s3_endpoint` | S3-compatible endpoint URL for Object Storage | No default |
| `flavor_name` | Instance size (tracking server runs on 2 vCPU / 2 GiB) | `s1a.small` |
| `image_name` | Operating system image | `Ubuntu-24.04` |
| `app_name` | Display name prefix for resources | `mlflow` |
| `volume_size` | Block volume size in GiB, mounted at `/var/lib/docker` | `20` |
| `external_network` | External network for floating IP allocation | `PublicStatic` |
| `private_cidr` | CIDR for the private subnet | `10.40.0.0/24` |
| `ui_allowed_cidr` | CIDR allowed to reach the UI on port 5000 | `10.40.0.0/24` |
| `postgres_db` | PostgreSQL database name | `mlflow` |
| `postgres_user` | PostgreSQL user | `mlflow` |
| `artifact_prefix` | Key prefix inside the artifact bucket | `mlflow` |

## UI access and security

The tracking UI listens on port 5000 over plain HTTP. The security group restricts 5000 to `ui_allowed_cidr`, which defaults to the private network only, so the raw UI stays off the public internet. Reach the UI one of three ways:

- Put a reverse proxy (Caddy or Nginx) in front of MLflow and serve the UI over HTTPS on 443. Point the domain's DNS A record at the floating IP. This is the recommended path for routine access.
- Tunnel over SSH: `ssh -L 5000:localhost:5000 ubuntu@FLOATING_IP`, then open `http://localhost:5000`.
- Set `ui_allowed_cidr` to `YOUR_IP/32` to reach port 5000 directly from one address.

Ports 80 and 443 stay open for the reverse proxy you put in front; they carry no traffic until you add one.


Experiment metadata lives in PostgreSQL and artifacts live in Object Storage. Snapshot the data volume before you resize or rebuild the host, and back up the external database on your Postgres instance.


## Metadata and artifacts

The tracking server splits storage by role:

- **PostgreSQL:** run metadata, parameters, metrics, and registry entries. Point `postgres_host`, `postgres_db`, and `postgres_user` at an external database, then add `POSTGRES_PASSWORD` to `/opt/mlflow/.env` on the instance and run `docker compose up -d --build`.
- **Object Storage:** models, plots, and other run artifacts. Set `artifact_bucket`, `artifact_prefix`, and `s3_endpoint`, then add `AWS_ACCESS_KEY_ID` and `AWS_SECRET_ACCESS_KEY` to `/opt/mlflow/.env`. Provision the bucket with the [S3 Storage with ACLs template](/resources/iac-templates/s3-storage-acl) or use an existing container.

Passwords and S3 credentials stay out of tfvars and the repo.

## CPU-only boundary

This template hosts the MLflow tracking server and model registry only. It does not provision GPUs, CUDA drivers, or training workers. Quake AI compute flavors are CPU-only AMD EPYC; experiment tracking fits that tier cleanly because logging metadata and artifacts is not training.

GPU training, fine-tuning, and heavy inference run on an external backend you operate. Point training jobs at that backend and set `MLFLOW_TRACKING_URI` to this server so runs still land in your registry. For GPU inference and RAG serving on Quake AI, see the [inference gateway template](/resources/iac-templates/inference-gateway).

## When to use this pattern

Run experiment tracking and a model registry on a VM you operate. MLflow suits notebook-driven exploration, batch training jobs that emit metrics, and teams that want a self-hosted alternative to SaaS experiment trackers.

For multi-user notebooks that log to MLflow, pair this template with [JupyterHub](/resources/iac-templates/jupyterhub). For the PostgreSQL metadata store, see [self-managed PostgreSQL](/resources/iac-templates/self-managed-postgres). For the artifact bucket, see [S3 Storage with ACLs](/resources/iac-templates/s3-storage-acl).

## Estimated cost

<PricingCompanion
  components={[
    { kind: "template", slug: "mlflow", required: true },
  ]}
/>

## Template source

<TemplateSource slug="mlflow" />

<TemplateResourceMap template="mlflow" format="opentofu" />

## Customize this pattern

- [Customize a template's image and flavor](/docs/automation/how-to/customize-template-image-flavor)
- [Add a block volume to a template](/docs/automation/how-to/add-volume-to-template)
- [Parameterize a template with a tfvars file](/docs/automation/how-to/parameterize-template-tfvars)

## See also

- [self-managed PostgreSQL](/resources/iac-templates/self-managed-postgres)
- [S3 Storage with ACLs](/resources/iac-templates/s3-storage-acl)
- [JupyterHub notebook server](/resources/iac-templates/jupyterhub)
- [inference gateway](/resources/iac-templates/inference-gateway)
