# Monitoring Stack (Prometheus + Grafana)

Source: https://docs.quake.ai/resources/iac-templates/monitoring-stack
Markdown: https://docs.quake.ai/resources/iac-templates/monitoring-stack.md

---

# Monitoring stack (Prometheus + Grafana)

This pattern composes Compute, Network, and Block Storage.

## What this template does

Provisions a self-hosted monitoring stack on Quake AI:

- Prometheus and Grafana instances configured through cloud-init user data ([cloud-init how-to](/docs/compute/how-to/configure-vm-cloud-init))
- Block storage volumes mounted at `/var/lib/prometheus` and `/var/lib/grafana` (on `/dev/sdb`) so metrics and dashboards persist on the dedicated volumes
- Private network for scrape traffic
- Floating IP on Grafana for dashboard access
- Pre-configured Prometheus scrape targets for Quake AI node exporters

The template requires a `grafana_admin_password` (no default), so the stack never ships with Grafana's `admin/admin`. By default the Grafana port and SSH are reachable from any address; set `admin_cidr` to your admin IP or CIDR to restrict access.

> **Security:** Grafana is exposed on the floating IP for dashboard access. Set a strong `grafana_admin_password` and restrict `admin_cidr` before you point real users or data at the stack.

## Parameters

| Parameter | Description | Default |
| --- | --- | --- |
| `key_name` | SSH keypair name (must already exist in your project) | required |
| `prometheus_flavor` | Prometheus instance size | `m2a.large` |
| `grafana_flavor` | Grafana instance size | `m2a.large` |
| `retention_days` | Prometheus data retention | `30` |
| `prometheus_volume_size` | Prometheus volume in GB | `100` |
| `grafana_volume_size` | Grafana volume in GB | `20` |
| `scrape_targets` | Initial scrape target list | `[]` |
| `grafana_admin_password` | Grafana admin password (required, no default) | _none_ |
| `admin_cidr` | Source CIDR allowed to reach Grafana and SSH | `0.0.0.0/0` |
| `image_name` | Operating system image | `Ubuntu-24.04` |
| `external_network` | External network for router gateway and floating IP | `PublicStatic` |
| `private_cidr` | Address range for the monitoring subnet | `192.168.90.0/24` |

## When to use this pattern

Deploy Prometheus and Grafana on dedicated instances with persistent block volumes. Pair this pattern with an application template such as [Three-Tier Application](/resources/iac-templates/three-tier-app) when you need metrics for separate web, application, and database tiers.

## Estimated cost

<PricingCompanion
  components={[
    { kind: "template", slug: "monitoring-stack", required: true },
  ]}
/>

## Template source

<TemplateSource slug="monitoring-stack" />

<TemplateResourceMap template="monitoring-stack" format="opentofu" />

## Customize this pattern

- [Customize a template's image and flavor](/docs/automation/how-to/customize-template-image-flavor)
- [Add a block volume to a template](/docs/automation/how-to/add-volume-to-template)
- [Parameterize a template with a tfvars file](/docs/automation/how-to/parameterize-template-tfvars)

## See also

- [Deploy the monitoring stack template with OpenTofu](/resources/deployments/deploy-monitoring-stack-template): end-to-end tutorial that applies this template, signs in to Grafana, confirms Prometheus scrape targets, and tears the stack down
- [Development Environment template](/resources/iac-templates/dev-environment)
- [Compute service overview](/docs/compute)
