Deploy MLflow with the mlflow template
Deploy MLflow with the mlflow template
Stand up MLflow, an open-source experiment-tracking and model-registry platform, on a single Quake AI instance using the validated OpenTofu template mlflow. You apply the template, wire it to an external PostgreSQL database and an Object Storage bucket, start the tracking server, and log a test run from your workstation.
MLflow is the tracking layer for ML platform teams. You run it yourself; this is a self-hosted tool you operate, not a managed service.
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on shared vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
MLflow tracking server
s1a.small · 2 shared vCPU, 2 GiB RAM, 0.5 Gbps
Runs the MLflow tracking server in Docker with PostgreSQL metadata and Object Storage artifacts.
The tracking server runs on 2 vCPU and 2 GiB RAM. This template hosts experiment tracking and the model registry only; GPU training runs elsewhere.
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
s1a.small
2 shared vCPU, 2 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
2 vCPU + 2 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (50 GiB)
50 GiB at $0.08/GiB/mo
Public IP (included)
1 included with the custom package
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Prerequisites#
You need:
- OpenTofu 1.6.0 or later (or Terraform 1.6.0 or later) installed locally.
- Your OpenStack credentials sourced into the shell (
source openrc.sh). See the OpenStack CLI guide. - An SSH keypair that already exists in your project. Record its name for the
key_namevariable. - A copy of the
mlflowtemplate directory from the template reference page. - An external PostgreSQL database with a database and user for MLflow metadata. The self-managed PostgreSQL template is one path; record the private IP for
postgres_host. - An Object Storage bucket and S3 credentials. The S3 Storage with ACLs template provisions both; record the bucket name, endpoint URL, access key, and secret key.
- Your workstation's public IP address, so you can open the UI port to it for first-boot setup. Find it with
curl -sS https://api.ipify.org.
A domain is optional for first boot. You add it in step 4 to serve the UI over HTTPS.
Step 1: Set the variables and apply the template#
The tracking UI listens on port 5000 over plain HTTP. The template's security group restricts port 5000 to ui_allowed_cidr, which defaults to the private network only, so the raw UI stays off the public internet. To reach the UI from your workstation for first-boot setup, set ui_allowed_cidr to your own address.
Copy the template's example variables file and open it:
cp terraform.tfvars.example terraform.tfvarsSet the required values:
key_name = "YOUR_KEY_NAME"
postgres_host = "POSTGRES_PRIVATE_IP"
artifact_bucket = "YOUR_BUCKET_NAME"
s3_endpoint = "https://us-east-1.rumble.cloud"
ui_allowed_cidr = "YOUR_IP/32"Initialize the working directory, preview the plan, and apply:
tofu init
tofu plan
tofu applyOpenTofu provisions a private network, a router, a security group, a block volume mounted at /var/lib/docker, an instance, and a floating IP. On first boot, cloud-init mounts the data volume and installs Docker Engine. MLflow does not start until you add credentials in step 2.
When the apply finishes, read the outputs:
tofu outputRecord floating_ip and tracking_url.
Step 2: Add credentials and start the tracking server#
No credential ships with this template. SSH to the instance and edit /opt/mlflow/.env. Uncomment and set the password and S3 credentials:
ssh ubuntu@YOUR_FLOATING_IP
sudo nano /opt/mlflow/.envAdd values for POSTGRES_PASSWORD, AWS_ACCESS_KEY_ID, and AWS_SECRET_ACCESS_KEY. The PostgreSQL user must already have access to the database named in POSTGRES_DB. Create the database on your Postgres instance first if it does not exist:
CREATE DATABASE mlflow;
CREATE USER mlflow WITH PASSWORD 'your-secure-password';
GRANT ALL PRIVILEGES ON DATABASE mlflow TO mlflow;Build the MLflow image and start the server:
cd /opt/mlflow
sudo docker compose up -d --build
sudo docker compose psOpen tracking_url (for example http://YOUR_FLOATING_IP:5000) in your browser. The MLflow home page loads when the container is healthy.
Step 3: Log a test experiment#
From your workstation, install the MLflow client and point it at the tracking server:
pip install mlflow
export MLFLOW_TRACKING_URI=http://YOUR_FLOATING_IP:5000Log a short run:
import mlflow
mlflow.set_experiment("quake-smoke-test")
with mlflow.start_run(run_name="hello-quake"):
mlflow.log_param("framework", "cpu-only")
mlflow.log_metric("accuracy", 0.91)
mlflow.set_tag("source", "deployment-walkthrough")Refresh the MLflow UI in your browser. The quake-smoke-test experiment appears with the run you logged.
This walkthrough logs metadata only. To log a model artifact, call mlflow.sklearn.log_model or mlflow.log_artifact in the same run; MLflow writes the file to the Object Storage bucket you configured.
Step 4: Serve the UI over HTTPS with Caddy#
The template leaves ports 80 and 443 open for a reverse proxy. Caddy
- Create a DNS A record for your domain (for example
mlflow.example.com) pointing atYOUR_FLOATING_IP. Follow How to point a domain at a Quake AI resource. Wait until the record resolves:
dig +short mlflow.example.com- SSH to the instance and create
/opt/mlflow/Caddyfile:
mlflow.example.com {
reverse_proxy 127.0.0.1:5000
}- Add Caddy to
/opt/mlflow/docker-compose.yml:
services:
caddy:
image: caddy:2
restart: unless-stopped
network_mode: host
volumes:
- /opt/mlflow/Caddyfile:/etc/caddy/Caddyfile
- caddy_data:/data
volumes:
caddy_data:- Apply the change:
cd /opt/mlflow
sudo docker compose up -dOpen https://mlflow.example.com and confirm the padlock. Once HTTPS works, close direct access to port 5000 by setting ui_allowed_cidr back to the private network in terraform.tfvars and running tofu apply.
Update training jobs and notebooks to use MLFLOW_TRACKING_URI=https://mlflow.example.com.
What you built#
- Applied the
mlflowtemplate to provision a network, security group, data volume, instance, and floating IP - Wired PostgreSQL and Object Storage by setting connection variables in tfvars and credentials in
/opt/mlflow/.env - Started the MLflow tracking server and logged a test experiment from your workstation
- Served the UI over HTTPS by pointing a domain at the floating IP and routing it through a Caddy reverse proxy
Scope of this deployment#
This template runs a single-VM MLflow tracking host, not a managed experiment-tracking cloud. The instance is CPU-only and runs in one region. You operate the instance, Docker, MLflow, the external PostgreSQL database, and the Object Storage bucket yourself: back them up, patch them, and watch storage growth as run volume increases.
This template hosts experiment tracking and the model registry only. GPU training and fine-tuning run on an external backend you operate. Point those jobs at your GPU cluster and set MLFLOW_TRACKING_URI to this server so metrics and artifacts still land in your registry.
Next steps#
- MLflow template reference: parameters, metadata and artifact wiring, and resource map
- self-managed PostgreSQL template: the metadata store this template points at
- S3 Storage with ACLs template: the artifact bucket this template points at
- Deploy JupyterHub with the jupyterhub template: multi-user notebooks that log runs to MLflow
- How to store application secrets and inject them at runtime: move database and S3 credentials out of plain environment files
Clean up#
When you no longer need the deployment, destroy everything the template created:
tofu destroyThe external PostgreSQL database and Object Storage bucket are not destroyed by this command; delete those resources separately if you no longer need them. Remove the DNS A record you created in step 4.
Quick answers
- Why does `openstack image save` write a 0-byte file for my boot-from-volume instance?CLI
- Why does `openstack server create` fail with "Only volume-backed servers are allowed for flavors with zero disk"?CLIAPITerraform
- Why does my project still have a 10 GiB Cinder volume after I deleted my instance?CLIAPI
See Also
Self-Managed PostgreSQL
Prerequisite
Migrate a Docker container app from AWS to Quake AI
Shares: Docker, Containers
Deploy Airbyte with the airbyte template
Shares: Docker, Containers
Deploy Airflow with the airflow template
Shares: Docker, Containers
Deploy Umami with the analytics-umami template
Shares: Docker, Containers