# Deploy the monitoring stack template with OpenTofu

Source: https://docs.quake.ai/resources/deployments/deploy-monitoring-stack-template
Markdown: https://docs.quake.ai/resources/deployments/deploy-monitoring-stack-template.md
> Stand up Prometheus and Grafana on a private subnet with Grafana reachable through one floating IP using the monitoring-stack OpenTofu template.

---

# Deploy the monitoring stack template with OpenTofu

Stand up Prometheus and Grafana on a private subnet using the [validated OpenTofu template](/docs/platform/validation#how-infrastructure-templates-are-checked) `monitoring-stack`. Prometheus stays on the private CIDR; Grafana gets one floating IP with a pre-provisioned Prometheus data source.

<PricingCompanion
  components={[
    { kind: "template", slug: "monitoring-stack", required: true },
  ]}
/>

<Figure size="md" caption="Monitoring stack topology: private subnet, Prometheus and Grafana instances with data volumes, one floating IP on Grafana">

```d2
direction: right

browser: Browser

cloud: Quake AI {
  fip: Floating IP\nGrafana :3000
  private: Private network\n192.168.90.0/24 {
    prom: Prometheus\n:9090 private only
    graf: Grafana\ndata source -> prom
  }
  router: Router\nto PublicStatic
}

browser -> cloud.fip: sign in
cloud.fip -> cloud.private.graf
cloud.private.graf -> cloud.private.prom: scrape
cloud.router -> cloud.private
```

</Figure>

## Prerequisites

You need:

- A Quake AI account with [application credentials](/docs/tools/generate-app-credentials)
- OpenTofu 1.6.0 or later ([installation guide](https://opentofu.org/docs/intro/install/))
- OpenStack credentials sourced into the shell (`source openrc.sh`). See [the OpenStack CLI guide](/docs/tools/openstack-cli).
- An SSH key pair already uploaded to the project. See [Add an SSH key](/docs/tools/add-ssh-key).
- A copy of the `monitoring-stack` template from [the template reference page](/resources/iac-templates/monitoring-stack)
- Enough project quota for two `m2a.large` instances (defaults), two 20 GB boot volumes, a 100 GB Prometheus data volume, a 20 GB Grafana data volume, one router, one private network, and one floating IP

## Step 1: Configure variables

Copy `terraform.tfvars.example` to `terraform.tfvars` and set:

```hcl
key_name                 = "YOUR_KEY_NAME"
grafana_admin_password   = "YOUR_STRONG_GRAFANA_PASSWORD"
admin_cidr               = "YOUR_PUBLIC_IP/32"
```

Replace `YOUR_PUBLIC_IP/32` with the address you browse from so Grafana port 3000 and SSH are not open to the entire internet.

Defaults for flavors, volume sizes, retention, and network CIDR are documented on the [Monitoring Stack (Prometheus + Grafana)](/resources/iac-templates/monitoring-stack) reference page.

## Step 2: Apply the template

From the template directory, run:

```bash
tofu init
tofu plan
tofu apply
```

Type `yes` when prompted. Provisioning takes several minutes while cloud-init installs Prometheus and Grafana on each instance.

When the run finishes, note `grafana_url` and `prometheus_private_ip` from the outputs.

## Step 3: Sign in to Grafana and confirm Prometheus

SSH to Grafana through its floating IP and wait for cloud-init:

```bash
GRAFANA_HOST=$(tofu output -raw grafana_url | sed 's|http://||;s|:3000||')
ssh -i ~/.ssh/YOUR_KEY -o StrictHostKeyChecking=accept-new ubuntu@"${GRAFANA_HOST}" 'cloud-init status --wait'
```

Open `grafana_url` in your browser. Sign in with username `admin` and the `grafana_admin_password` value from `terraform.tfvars`.

Grafana ships with a provisioned Prometheus data source pointing at `http://<prometheus_private_ip>:9090`. Go to **Connections** > **Data sources** > **Prometheus** and select **Save & test**.

Check scrape targets from the Grafana instance:

```bash
PROM_IP=$(tofu output -raw prometheus_private_ip)
ssh -i ~/.ssh/YOUR_KEY ubuntu@"${GRAFANA_HOST}" \
  "curl -s http://${PROM_IP}:9090/api/v1/targets | python3 -m json.tool | head -40"
```

You should see at least the `prometheus` job scraping `localhost:9090` with state `up`.

In Grafana, open **Explore**, select the Prometheus data source, and run `up`. A result with `up{job="prometheus"} = 1` confirms metrics are flowing.

## Next steps

- [Monitor your Quake AI workload with Prometheus and Grafana](/docs/operate/monitoring/how-to-deploy-monitoring)
- [Ship application logs off your VMs](/docs/operate/monitoring/how-to-ship-logs)
- [Monitoring Stack template](/resources/iac-templates/monitoring-stack)
- [Three-Tier Application template](/resources/iac-templates/three-tier-app)

## Clean up

Run `tofu destroy` from the project directory when finished. Type `yes` to confirm. Verify in the Console that both instances and the floating IP are gone.
