Monitoring Stack (Prometheus + Grafana)
Monitoring stack (Prometheus + Grafana)
This pattern composes Compute, Network, and Block Storage.
What this template does#
Provisions a self-hosted monitoring stack on Quake AI:
- Prometheus and Grafana instances configured through cloud-init user data (cloud-init how-to)
- Block storage volumes mounted at
/var/lib/prometheusand/var/lib/grafana(on/dev/sdb) so metrics and dashboards persist on the dedicated volumes - Private network for scrape traffic
- Floating IP on Grafana for dashboard access
- Pre-configured Prometheus scrape targets for Quake AI node exporters
The template requires a grafana_admin_password (no default), so the stack never ships with Grafana's admin/admin. By default the Grafana port and SSH are reachable from any address; set admin_cidr to your admin IP or CIDR to restrict access.
Security: Grafana is exposed on the floating IP for dashboard access. Set a strong
grafana_admin_passwordand restrictadmin_cidrbefore you point real users or data at the stack.
Parameters#
| Parameter | Description | Default |
|---|---|---|
key_name | SSH keypair name (must already exist in your project) | required |
prometheus_flavor | Prometheus instance size | m2a.large |
grafana_flavor | Grafana instance size | m2a.large |
retention_days | Prometheus data retention | 30 |
prometheus_volume_size | Prometheus volume in GB | 100 |
grafana_volume_size | Grafana volume in GB | 20 |
scrape_targets | Initial scrape target list | [] |
grafana_admin_password | Grafana admin password (required, no default) | none |
admin_cidr | Source CIDR allowed to reach Grafana and SSH | 0.0.0.0/0 |
image_name | Operating system image | Ubuntu-24.04 |
external_network | External network for router gateway and floating IP | PublicStatic |
private_cidr | Address range for the monitoring subnet | 192.168.90.0/24 |
When to use this pattern#
Deploy Prometheus and Grafana on dedicated instances with persistent block volumes. Pair this pattern with an application template such as Three-Tier Application when you need metrics for separate web, application, and database tiers.
Estimated cost#
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Prometheus
m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Grafana
m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
m2a.large
2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
m2a.large
2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
4 vCPU + 16 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (160 GiB)
160 GiB at $0.08/GiB/mo
Public IP (included)
1 included with the custom package
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on dedicated CPU
The headline estimate above; predictable steady-load performance.
Saves $99.00/mo while you build on shared CPU.
Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM); m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Template source#
Show source (8 files)Hide source
locals {
scrape_targets_yaml = length(var.scrape_targets) == 0 ? "" : join("\n", concat(
[
" - job_name: \"external\"",
" static_configs:",
" - targets:"
],
[for t in var.scrape_targets : " - ${yamlencode(t)}"]
))
}
data "openstack_images_image_v2" "os" {
name = var.image_name
most_recent = true
}
data "openstack_networking_network_v2" "external" {
name = var.external_network
}
resource "openstack_networking_network_v2" "private" {
name = "monitoring-private"
admin_state_up = true
}
resource "openstack_networking_subnet_v2" "private" {
name = "monitoring-private-sn"
network_id = openstack_networking_network_v2.private.id
cidr = var.private_cidr
ip_version = 4
enable_dhcp = true
dns_nameservers = ["8.8.8.8", "8.8.4.4"]
}
resource "openstack_networking_router_v2" "router" {
name = "monitoring-router"
external_network_id = data.openstack_networking_network_v2.external.id
admin_state_up = true
}
resource "openstack_networking_router_interface_v2" "private" {
router_id = openstack_networking_router_v2.router.id
subnet_id = openstack_networking_subnet_v2.private.id
}
resource "openstack_networking_secgroup_v2" "prometheus" {
name = "monitoring-prometheus"
description = "Prometheus 9090 from private CIDR; SSH from anywhere"
}
resource "openstack_networking_secgroup_rule_v2" "prometheus_http" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 9090
port_range_max = 9090
remote_ip_prefix = var.private_cidr
security_group_id = openstack_networking_secgroup_v2.prometheus.id
}
resource "openstack_networking_secgroup_rule_v2" "prometheus_ssh" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 22
port_range_max = 22
remote_ip_prefix = var.admin_cidr
security_group_id = openstack_networking_secgroup_v2.prometheus.id
}
resource "openstack_networking_secgroup_v2" "grafana" {
name = "monitoring-grafana"
description = "Grafana 3000 from anywhere; SSH from anywhere"
}
resource "openstack_networking_secgroup_rule_v2" "grafana_http" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 3000
port_range_max = 3000
remote_ip_prefix = var.admin_cidr
security_group_id = openstack_networking_secgroup_v2.grafana.id
}
resource "openstack_networking_secgroup_rule_v2" "grafana_ssh" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 22
port_range_max = 22
remote_ip_prefix = var.admin_cidr
security_group_id = openstack_networking_secgroup_v2.grafana.id
}
resource "openstack_blockstorage_volume_v3" "prometheus_data" {
name = "monitoring-prometheus-data"
size = var.prometheus_volume_size
}
resource "openstack_blockstorage_volume_v3" "grafana_data" {
name = "monitoring-grafana-data"
size = var.grafana_volume_size
}
resource "openstack_networking_port_v2" "prometheus" {
name = "monitoring-prometheus-port"
network_id = openstack_networking_network_v2.private.id
security_group_ids = [openstack_networking_secgroup_v2.prometheus.id]
fixed_ip {
subnet_id = openstack_networking_subnet_v2.private.id
}
depends_on = [openstack_networking_router_interface_v2.private]
}
resource "openstack_compute_instance_v2" "prometheus" {
name = "monitoring-prometheus"
flavor_name = var.prometheus_flavor
key_pair = var.key_name
user_data = templatefile("${path.module}/cloud-init/prometheus.yaml", {
retention_days = var.retention_days
scrape_targets_yaml = local.scrape_targets_yaml
})
block_device {
uuid = data.openstack_images_image_v2.os.id
source_type = "image"
destination_type = "volume"
volume_size = 20
boot_index = 0
delete_on_termination = true
}
block_device {
uuid = openstack_blockstorage_volume_v3.prometheus_data.id
source_type = "volume"
destination_type = "volume"
boot_index = -1
delete_on_termination = false
}
network {
port = openstack_networking_port_v2.prometheus.id
}
}
resource "openstack_compute_instance_v2" "grafana" {
name = "monitoring-grafana"
flavor_name = var.grafana_flavor
key_pair = var.key_name
user_data = templatefile("${path.module}/cloud-init/grafana.yaml", {
prometheus_ip = openstack_compute_instance_v2.prometheus.network[0].fixed_ip_v4
grafana_admin_password = var.grafana_admin_password
})
block_device {
uuid = data.openstack_images_image_v2.os.id
source_type = "image"
destination_type = "volume"
volume_size = 20
boot_index = 0
delete_on_termination = true
}
block_device {
uuid = openstack_blockstorage_volume_v3.grafana_data.id
source_type = "volume"
destination_type = "volume"
boot_index = -1
delete_on_termination = false
}
network {
port = openstack_networking_port_v2.grafana.id
}
depends_on = [
openstack_compute_instance_v2.prometheus,
]
}
resource "openstack_networking_port_v2" "grafana" {
name = "monitoring-grafana-port"
network_id = openstack_networking_network_v2.private.id
security_group_ids = [openstack_networking_secgroup_v2.grafana.id]
fixed_ip {
subnet_id = openstack_networking_subnet_v2.private.id
}
depends_on = [openstack_networking_router_interface_v2.private]
}
resource "openstack_networking_floatingip_v2" "grafana" {
pool = var.external_network
}
resource "openstack_networking_floatingip_associate_v2" "grafana" {
floating_ip = openstack_networking_floatingip_v2.grafana.address
port_id = openstack_networking_port_v2.grafana.id
}
variable "key_name" {
description = "Existing SSH keypair name in the project for compute instances"
type = string
}
variable "prometheus_flavor" {
description = "Flavor for the Prometheus instance"
type = string
default = "m2a.large"
}
variable "grafana_flavor" {
description = "Flavor for the Grafana instance"
type = string
default = "m2a.large"
}
variable "retention_days" {
description = "Prometheus TSDB retention in days"
type = number
default = 30
}
variable "prometheus_volume_size" {
description = "Cinder volume size in GiB for Prometheus TSDB"
type = number
default = 100
}
variable "grafana_volume_size" {
description = "Cinder volume size in GiB for Grafana data"
type = number
default = 20
}
variable "scrape_targets" {
description = "Additional scrape targets (host:port)"
type = list(string)
default = []
}
variable "image_name" {
description = "Boot image name"
type = string
default = "Ubuntu-24.04"
}
variable "external_network" {
description = "Shared external network for router gateway and floating IPs; defaults to PublicStatic (persisted FIP / production pattern). Override with PublicEphemeral for ephemeral demos."
type = string
default = "PublicStatic"
}
variable "private_cidr" {
description = "CIDR for the private subnet"
type = string
default = "192.168.90.0/24"
}
variable "admin_cidr" {
description = "Source CIDR allowed to reach Grafana (3000) and SSH (22). Restrict this to your admin IP/CIDR; the default leaves Grafana reachable from any address."
type = string
default = "0.0.0.0/0"
}
variable "grafana_admin_password" {
description = "Grafana admin password. Required, no default, so the stack never ships with admin/admin. Set a strong value when you apply."
type = string
sensitive = true
}
output "grafana_url" {
description = "Grafana HTTP URL on the floating IP"
value = "http://${openstack_networking_floatingip_v2.grafana.address}:3000"
}
output "prometheus_private_ip" {
description = "Private IPv4 of the Prometheus instance"
value = openstack_compute_instance_v2.prometheus.network[0].fixed_ip_v4
}
terraform {
required_version = ">= 1.6.0"
required_providers {
openstack = {
source = "terraform-provider-openstack/openstack"
version = "~> 2.0"
}
}
}
provider "openstack" {}
# Required
key_name = "YOUR_KEY_NAME"
# prometheus_flavor = "m2a.large"
# grafana_flavor = "m2a.large"
# retention_days = 30
# prometheus_volume_size = 100
# grafana_volume_size = 20
# scrape_targets = []
# image_name = "Ubuntu-24.04"
# external_network = "PublicStatic"
# private_cidr = "192.168.90.0/24"
#cloud-config
package_update: true
packages:
- apt-transport-https
- ca-certificates
- curl
- gnupg
write_files:
- path: /etc/apt/sources.list.d/grafana.list
content: deb [signed-by=/usr/share/keyrings/grafana.gpg] https://apt.grafana.com stable main
owner: root:root
permissions: "0644"
- path: /etc/grafana/provisioning/datasources/prometheus.yaml
content: |
apiVersion: 1
datasources:
- name: Prometheus
type: prometheus
access: proxy
url: http://${prometheus_ip}:9090
isDefault: true
editable: false
owner: root:root
permissions: "0644"
- path: /etc/systemd/system/grafana-server.service.d/override.conf
content: |
[Service]
Environment=GF_SECURITY_ADMIN_PASSWORD=${grafana_admin_password}
owner: root:root
permissions: "0600"
runcmd:
- |
set -e
export DEBIAN_FRONTEND=noninteractive
# The data volume attaches as /dev/sdb on this platform (not /dev/vdb).
DEV=/dev/sdb
for i in $(seq 1 30); do [ -b "$DEV" ] && break; sleep 5; done
if ! blkid "$DEV" >/dev/null 2>&1; then mkfs.ext4 -F -L grafdata "$DEV"; fi
mkdir -p /var/lib/grafana
mount "$DEV" /var/lib/grafana
grep -q "$DEV" /etc/fstab || echo "$DEV /var/lib/grafana ext4 defaults,nofail 0 2" >> /etc/fstab
install -d /usr/share/keyrings
curl -fsSL https://apt.grafana.com/gpg.key | gpg --batch --yes --dearmor -o /usr/share/keyrings/grafana.gpg
for i in 1 2 3; do apt-get update && break; sleep 10; done
apt-get install -y grafana
chown -R grafana:grafana /var/lib/grafana
systemctl daemon-reload
systemctl enable grafana-server
systemctl restart grafana-server
#cloud-config
package_update: true
write_files:
- path: /tmp/prometheus.default
content: |
ARGS="--web.enable-lifecycle --storage.tsdb.retention.time=${retention_days}d"
owner: root:root
permissions: "0644"
- path: /tmp/prometheus.yml
content: |
global:
scrape_interval: 15s
scrape_configs:
- job_name: "prometheus"
static_configs:
- targets: ["localhost:9090"]
${scrape_targets_yaml}
owner: root:root
permissions: "0644"
runcmd:
- |
set -e
export DEBIAN_FRONTEND=noninteractive
# The data volume attaches as /dev/sdb on this platform (not /dev/vdb).
DEV=/dev/sdb
for i in $(seq 1 30); do [ -b "$DEV" ] && break; sleep 5; done
if ! blkid "$DEV" >/dev/null 2>&1; then mkfs.ext4 -F -L promdata "$DEV"; fi
mkdir -p /var/lib/prometheus
mount "$DEV" /var/lib/prometheus
grep -q "$DEV" /etc/fstab || echo "$DEV /var/lib/prometheus ext4 defaults,nofail 0 2" >> /etc/fstab
for i in 1 2 3; do apt-get update && break; sleep 10; done
apt-get install -y prometheus
cp /tmp/prometheus.yml /etc/prometheus/prometheus.yml
cp /tmp/prometheus.default /etc/default/prometheus
chown -R prometheus:prometheus /var/lib/prometheus
systemctl enable prometheus
systemctl restart prometheus
# Monitoring Stack
Prometheus and Grafana on dedicated instances with persistent volumes, configurable retention, and optional extra scrape targets.
**Network class:** production — `external_network` defaults to `PublicStatic` for persisted floating IPs and multi-tier stacks; override with `PublicEphemeral` for ephemeral demos.
## Prerequisites
- OpenTofu >= 1.6.0 or Terraform >= 1.6.0
- Quake AI account with OpenStack credentials
- Quota for two compute instances and two block volumes (sizes configurable)
## Usage
1. Clone or copy this template directory
2. Copy `terraform.tfvars.example` to `terraform.tfvars` and fill in your values
3. Source your OpenStack credentials: `source openrc.sh`
4. Initialize: `tofu init`
5. Preview: `tofu plan`
6. Apply: `tofu apply`
## Variables
| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `key_name` | string | yes | n/a | Existing SSH keypair name in the project for compute instances |
| `prometheus_flavor` | string | no | `m2a.large` | Flavor for the Prometheus instance |
| `grafana_flavor` | string | no | `m2a.large` | Flavor for the Grafana instance |
| `retention_days` | number | no | `30` | Prometheus TSDB retention in days |
| `prometheus_volume_size` | number | no | `100` | Cinder volume size in GiB for Prometheus TSDB |
| `grafana_volume_size` | number | no | `20` | Cinder volume size in GiB for Grafana data |
| `scrape_targets` | list(string) | no | `[]` | Additional scrape targets (`host:port`) |
| `image_name` | string | no | `Ubuntu-24.04` | Boot image name |
| `external_network` | string | no | `PublicStatic` | Persisted FIP / production default; override with `PublicEphemeral` for demos |
| `private_cidr` | string | no | `192.168.90.0/24` | Private subnet CIDR |
## Documentation
Full documentation: [Monitoring stack template](/docs/automation/templates/monitoring-stack)
Resources, parameters, and variables
key_namerequiredprometheus_flavor="m2a.large"grafana_flavor="m2a.large"retention_days=30prometheus_volume_size=100grafana_volume_size=20scrape_targets=[]image_name="Ubuntu-24.04"external_network="PublicStatic"private_cidr="192.168.90.0/24"admin_cidr="0.0.0.0/0"grafana_admin_passwordrequired
Customize this pattern#
- Customize a template's image and flavor
- Add a block volume to a template
- Parameterize a template with a tfvars file
See also#
- Deploy the monitoring stack template with OpenTofu: end-to-end tutorial that applies this template, signs in to Grafana, confirms Prometheus scrape targets, and tears the stack down
- Development Environment template
- Compute service overview
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.
Last validated: 07.07.2026
See Also
Terraform and OpenTofu on Quake AI
Prerequisite
Networks
Prerequisite
Authoring IaC templates for Quake AI
Shares: Volumes, Security Groups
Deploy an API gateway with the api-gateway template
Shares: Volumes, Security Groups
Deploy a regional edge cache with the edge-cache template
Shares: Volumes, Security Groups