Skip to content
IaC Templates

Monitoring Stack (Prometheus + Grafana)

Template · Updated Jul 2026
Validated Jul 2026

Monitoring stack (Prometheus + Grafana)

This pattern composes Compute, Network, and Block Storage.

What this template does#

Provisions a self-hosted monitoring stack on Quake AI:

  • Prometheus and Grafana instances configured through cloud-init user data (cloud-init how-to)
  • Block storage volumes mounted at /var/lib/prometheus and /var/lib/grafana (on /dev/sdb) so metrics and dashboards persist on the dedicated volumes
  • Private network for scrape traffic
  • Floating IP on Grafana for dashboard access
  • Pre-configured Prometheus scrape targets for Quake AI node exporters

The template requires a grafana_admin_password (no default), so the stack never ships with Grafana's admin/admin. By default the Grafana port and SSH are reachable from any address; set admin_cidr to your admin IP or CIDR to restrict access.

Security: Grafana is exposed on the floating IP for dashboard access. Set a strong grafana_admin_password and restrict admin_cidr before you point real users or data at the stack.

Parameters#

ParameterDescriptionDefault
key_nameSSH keypair name (must already exist in your project)required
prometheus_flavorPrometheus instance sizem2a.large
grafana_flavorGrafana instance sizem2a.large
retention_daysPrometheus data retention30
prometheus_volume_sizePrometheus volume in GB100
grafana_volume_sizeGrafana volume in GB20
scrape_targetsInitial scrape target list[]
grafana_admin_passwordGrafana admin password (required, no default)none
admin_cidrSource CIDR allowed to reach Grafana and SSH0.0.0.0/0
image_nameOperating system imageUbuntu-24.04
external_networkExternal network for router gateway and floating IPPublicStatic
private_cidrAddress range for the monitoring subnet192.168.90.0/24

When to use this pattern#

Deploy Prometheus and Grafana on dedicated instances with persistent block volumes. Pair this pattern with an application template such as Three-Tier Application when you need metrics for separate web, application, and database tiers.

Estimated cost#

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on dedicated vCPU.

Starting template$139.80/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Prometheus

m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00/mo

Grafana

m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

m2a.large

2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00

m2a.large

2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00

Compute + RAM rate basis

4 vCPU + 16 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (160 GiB)

160 GiB at $0.08/GiB/mo

$12.80

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$40.80/mo

Production on dedicated CPU

The headline estimate above; predictable steady-load performance.

$139.80/mo

Saves $99.00/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM); m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Template source#

8 files. Download the zip or expand to copy any file.Download monitoring-stack.zip
Show source (8 files)
main.tfHCL
locals {
  scrape_targets_yaml = length(var.scrape_targets) == 0 ? "" : join("\n", concat(
    [
      "        - job_name: \"external\"",
      "          static_configs:",
      "            - targets:"
    ],
    [for t in var.scrape_targets : "                - ${yamlencode(t)}"]
  ))
}

data "openstack_images_image_v2" "os" {
  name        = var.image_name
  most_recent = true
}

data "openstack_networking_network_v2" "external" {
  name = var.external_network
}


resource "openstack_networking_network_v2" "private" {
  name           = "monitoring-private"
  admin_state_up = true
}

resource "openstack_networking_subnet_v2" "private" {
  name            = "monitoring-private-sn"
  network_id      = openstack_networking_network_v2.private.id
  cidr            = var.private_cidr
  ip_version      = 4
  enable_dhcp     = true
  dns_nameservers = ["8.8.8.8", "8.8.4.4"]
}

resource "openstack_networking_router_v2" "router" {
  name                = "monitoring-router"
  external_network_id = data.openstack_networking_network_v2.external.id
  admin_state_up      = true
}

resource "openstack_networking_router_interface_v2" "private" {
  router_id = openstack_networking_router_v2.router.id
  subnet_id = openstack_networking_subnet_v2.private.id
}

resource "openstack_networking_secgroup_v2" "prometheus" {
  name        = "monitoring-prometheus"
  description = "Prometheus 9090 from private CIDR; SSH from anywhere"
}

resource "openstack_networking_secgroup_rule_v2" "prometheus_http" {
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 9090
  port_range_max    = 9090
  remote_ip_prefix  = var.private_cidr
  security_group_id = openstack_networking_secgroup_v2.prometheus.id
}

resource "openstack_networking_secgroup_rule_v2" "prometheus_ssh" {
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 22
  port_range_max    = 22
  remote_ip_prefix  = var.admin_cidr
  security_group_id = openstack_networking_secgroup_v2.prometheus.id
}

resource "openstack_networking_secgroup_v2" "grafana" {
  name        = "monitoring-grafana"
  description = "Grafana 3000 from anywhere; SSH from anywhere"
}

resource "openstack_networking_secgroup_rule_v2" "grafana_http" {
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 3000
  port_range_max    = 3000
  remote_ip_prefix  = var.admin_cidr
  security_group_id = openstack_networking_secgroup_v2.grafana.id
}

resource "openstack_networking_secgroup_rule_v2" "grafana_ssh" {
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 22
  port_range_max    = 22
  remote_ip_prefix  = var.admin_cidr
  security_group_id = openstack_networking_secgroup_v2.grafana.id
}

resource "openstack_blockstorage_volume_v3" "prometheus_data" {
  name = "monitoring-prometheus-data"
  size = var.prometheus_volume_size
}

resource "openstack_blockstorage_volume_v3" "grafana_data" {
  name = "monitoring-grafana-data"
  size = var.grafana_volume_size
}

resource "openstack_networking_port_v2" "prometheus" {
  name               = "monitoring-prometheus-port"
  network_id         = openstack_networking_network_v2.private.id
  security_group_ids = [openstack_networking_secgroup_v2.prometheus.id]

  fixed_ip {
    subnet_id = openstack_networking_subnet_v2.private.id
  }

  depends_on = [openstack_networking_router_interface_v2.private]
}

resource "openstack_compute_instance_v2" "prometheus" {
  name        = "monitoring-prometheus"
  flavor_name = var.prometheus_flavor
  key_pair    = var.key_name

  user_data = templatefile("${path.module}/cloud-init/prometheus.yaml", {
    retention_days      = var.retention_days
    scrape_targets_yaml = local.scrape_targets_yaml
  })

  block_device {
    uuid                  = data.openstack_images_image_v2.os.id
    source_type           = "image"
    destination_type      = "volume"
    volume_size           = 20
    boot_index            = 0
    delete_on_termination = true
  }

  block_device {
    uuid                  = openstack_blockstorage_volume_v3.prometheus_data.id
    source_type           = "volume"
    destination_type      = "volume"
    boot_index            = -1
    delete_on_termination = false
  }

  network {
    port = openstack_networking_port_v2.prometheus.id
  }
}

resource "openstack_compute_instance_v2" "grafana" {
  name        = "monitoring-grafana"
  flavor_name = var.grafana_flavor
  key_pair    = var.key_name

  user_data = templatefile("${path.module}/cloud-init/grafana.yaml", {
    prometheus_ip          = openstack_compute_instance_v2.prometheus.network[0].fixed_ip_v4
    grafana_admin_password = var.grafana_admin_password
  })

  block_device {
    uuid                  = data.openstack_images_image_v2.os.id
    source_type           = "image"
    destination_type      = "volume"
    volume_size           = 20
    boot_index            = 0
    delete_on_termination = true
  }

  block_device {
    uuid                  = openstack_blockstorage_volume_v3.grafana_data.id
    source_type           = "volume"
    destination_type      = "volume"
    boot_index            = -1
    delete_on_termination = false
  }

  network {
    port = openstack_networking_port_v2.grafana.id
  }

  depends_on = [
    openstack_compute_instance_v2.prometheus,
  ]
}

resource "openstack_networking_port_v2" "grafana" {
  name               = "monitoring-grafana-port"
  network_id         = openstack_networking_network_v2.private.id
  security_group_ids = [openstack_networking_secgroup_v2.grafana.id]

  fixed_ip {
    subnet_id = openstack_networking_subnet_v2.private.id
  }

  depends_on = [openstack_networking_router_interface_v2.private]
}

resource "openstack_networking_floatingip_v2" "grafana" {
  pool = var.external_network
}

resource "openstack_networking_floatingip_associate_v2" "grafana" {
  floating_ip = openstack_networking_floatingip_v2.grafana.address
  port_id     = openstack_networking_port_v2.grafana.id
}
variables.tfHCL
variable "key_name" {
  description = "Existing SSH keypair name in the project for compute instances"
  type        = string
}

variable "prometheus_flavor" {
  description = "Flavor for the Prometheus instance"
  type        = string
  default     = "m2a.large"
}

variable "grafana_flavor" {
  description = "Flavor for the Grafana instance"
  type        = string
  default     = "m2a.large"
}

variable "retention_days" {
  description = "Prometheus TSDB retention in days"
  type        = number
  default     = 30
}

variable "prometheus_volume_size" {
  description = "Cinder volume size in GiB for Prometheus TSDB"
  type        = number
  default     = 100
}

variable "grafana_volume_size" {
  description = "Cinder volume size in GiB for Grafana data"
  type        = number
  default     = 20
}

variable "scrape_targets" {
  description = "Additional scrape targets (host:port)"
  type        = list(string)
  default     = []
}

variable "image_name" {
  description = "Boot image name"
  type        = string
  default     = "Ubuntu-24.04"
}

variable "external_network" {
  description = "Shared external network for router gateway and floating IPs; defaults to PublicStatic (persisted FIP / production pattern). Override with PublicEphemeral for ephemeral demos."
  type        = string
  default     = "PublicStatic"
}

variable "private_cidr" {
  description = "CIDR for the private subnet"
  type        = string
  default     = "192.168.90.0/24"
}

variable "admin_cidr" {
  description = "Source CIDR allowed to reach Grafana (3000) and SSH (22). Restrict this to your admin IP/CIDR; the default leaves Grafana reachable from any address."
  type        = string
  default     = "0.0.0.0/0"
}

variable "grafana_admin_password" {
  description = "Grafana admin password. Required, no default, so the stack never ships with admin/admin. Set a strong value when you apply."
  type        = string
  sensitive   = true
}
outputs.tfHCL
output "grafana_url" {
  description = "Grafana HTTP URL on the floating IP"
  value       = "http://${openstack_networking_floatingip_v2.grafana.address}:3000"
}

output "prometheus_private_ip" {
  description = "Private IPv4 of the Prometheus instance"
  value       = openstack_compute_instance_v2.prometheus.network[0].fixed_ip_v4
}
versions.tfHCL
terraform {
  required_version = ">= 1.6.0"

  required_providers {
    openstack = {
      source  = "terraform-provider-openstack/openstack"
      version = "~> 2.0"
    }
  }
}

provider "openstack" {}
terraform.tfvars.exampleHCL
# Required
key_name = "YOUR_KEY_NAME"

# prometheus_flavor = "m2a.large"
# grafana_flavor = "m2a.large"
# retention_days = 30
# prometheus_volume_size = 100
# grafana_volume_size = 20
# scrape_targets = []
# image_name = "Ubuntu-24.04"
# external_network = "PublicStatic"
# private_cidr = "192.168.90.0/24"
cloud-init/grafana.yamlYAML
#cloud-config
package_update: true
packages:
  - apt-transport-https
  - ca-certificates
  - curl
  - gnupg

write_files:
  - path: /etc/apt/sources.list.d/grafana.list
    content: deb [signed-by=/usr/share/keyrings/grafana.gpg] https://apt.grafana.com stable main
    owner: root:root
    permissions: "0644"
  - path: /etc/grafana/provisioning/datasources/prometheus.yaml
    content: |
      apiVersion: 1
      datasources:
        - name: Prometheus
          type: prometheus
          access: proxy
          url: http://${prometheus_ip}:9090
          isDefault: true
          editable: false
    owner: root:root
    permissions: "0644"
  - path: /etc/systemd/system/grafana-server.service.d/override.conf
    content: |
      [Service]
      Environment=GF_SECURITY_ADMIN_PASSWORD=${grafana_admin_password}
    owner: root:root
    permissions: "0600"

runcmd:
  - |
    set -e
    export DEBIAN_FRONTEND=noninteractive
    # The data volume attaches as /dev/sdb on this platform (not /dev/vdb).
    DEV=/dev/sdb
    for i in $(seq 1 30); do [ -b "$DEV" ] && break; sleep 5; done
    if ! blkid "$DEV" >/dev/null 2>&1; then mkfs.ext4 -F -L grafdata "$DEV"; fi
    mkdir -p /var/lib/grafana
    mount "$DEV" /var/lib/grafana
    grep -q "$DEV" /etc/fstab || echo "$DEV /var/lib/grafana ext4 defaults,nofail 0 2" >> /etc/fstab
    install -d /usr/share/keyrings
    curl -fsSL https://apt.grafana.com/gpg.key | gpg --batch --yes --dearmor -o /usr/share/keyrings/grafana.gpg
    for i in 1 2 3; do apt-get update && break; sleep 10; done
    apt-get install -y grafana
    chown -R grafana:grafana /var/lib/grafana
    systemctl daemon-reload
    systemctl enable grafana-server
    systemctl restart grafana-server
cloud-init/prometheus.yamlYAML
#cloud-config
package_update: true
write_files:
  - path: /tmp/prometheus.default
    content: |
      ARGS="--web.enable-lifecycle --storage.tsdb.retention.time=${retention_days}d"
    owner: root:root
    permissions: "0644"
  - path: /tmp/prometheus.yml
    content: |
      global:
        scrape_interval: 15s
      scrape_configs:
        - job_name: "prometheus"
          static_configs:
            - targets: ["localhost:9090"]
${scrape_targets_yaml}
    owner: root:root
    permissions: "0644"
runcmd:
  - |
    set -e
    export DEBIAN_FRONTEND=noninteractive
    # The data volume attaches as /dev/sdb on this platform (not /dev/vdb).
    DEV=/dev/sdb
    for i in $(seq 1 30); do [ -b "$DEV" ] && break; sleep 5; done
    if ! blkid "$DEV" >/dev/null 2>&1; then mkfs.ext4 -F -L promdata "$DEV"; fi
    mkdir -p /var/lib/prometheus
    mount "$DEV" /var/lib/prometheus
    grep -q "$DEV" /etc/fstab || echo "$DEV /var/lib/prometheus ext4 defaults,nofail 0 2" >> /etc/fstab
    for i in 1 2 3; do apt-get update && break; sleep 10; done
    apt-get install -y prometheus
    cp /tmp/prometheus.yml /etc/prometheus/prometheus.yml
    cp /tmp/prometheus.default /etc/default/prometheus
    chown -R prometheus:prometheus /var/lib/prometheus
    systemctl enable prometheus
    systemctl restart prometheus
README.mdMarkdown
# Monitoring Stack

Prometheus and Grafana on dedicated instances with persistent volumes, configurable retention, and optional extra scrape targets.


**Network class:** production — `external_network` defaults to `PublicStatic` for persisted floating IPs and multi-tier stacks; override with `PublicEphemeral` for ephemeral demos.

## Prerequisites

- OpenTofu >= 1.6.0 or Terraform >= 1.6.0
- Quake AI account with OpenStack credentials
- Quota for two compute instances and two block volumes (sizes configurable)

## Usage

1. Clone or copy this template directory
2. Copy `terraform.tfvars.example` to `terraform.tfvars` and fill in your values
3. Source your OpenStack credentials: `source openrc.sh`
4. Initialize: `tofu init`
5. Preview: `tofu plan`
6. Apply: `tofu apply`

## Variables

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `key_name` | string | yes | n/a | Existing SSH keypair name in the project for compute instances |
| `prometheus_flavor` | string | no | `m2a.large` | Flavor for the Prometheus instance |
| `grafana_flavor` | string | no | `m2a.large` | Flavor for the Grafana instance |
| `retention_days` | number | no | `30` | Prometheus TSDB retention in days |
| `prometheus_volume_size` | number | no | `100` | Cinder volume size in GiB for Prometheus TSDB |
| `grafana_volume_size` | number | no | `20` | Cinder volume size in GiB for Grafana data |
| `scrape_targets` | list(string) | no | `[]` | Additional scrape targets (`host:port`) |
| `image_name` | string | no | `Ubuntu-24.04` | Boot image name |
| `external_network` | string | no | `PublicStatic` | Persisted FIP / production default; override with `PublicEphemeral` for demos |
| `private_cidr` | string | no | `192.168.90.0/24` | Private subnet CIDR |

## Documentation

Full documentation: [Monitoring stack template](/docs/automation/templates/monitoring-stack)
Resources, parameters, and variables
Provisions
Parameterized by
Variables
  • key_namerequired
  • prometheus_flavor="m2a.large"
  • grafana_flavor="m2a.large"
  • retention_days=30
  • prometheus_volume_size=100
  • grafana_volume_size=20
  • scrape_targets=[]
  • image_name="Ubuntu-24.04"
  • external_network="PublicStatic"
  • private_cidr="192.168.90.0/24"
  • admin_cidr="0.0.0.0/0"
  • grafana_admin_passwordrequired

Customize this pattern#

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

For the full policy, see Usage Guidelines.

Last validated: 07.07.2026

Was this page helpful?