# How to monitor your Quake AI workload with Prometheus and Grafana

Source: https://docs.quake.ai/docs/operate/monitoring/how-to-deploy-monitoring
Markdown: https://docs.quake.ai/docs/operate/monitoring/how-to-deploy-monitoring.md

---

# How to monitor your Quake AI workload with Prometheus and Grafana

Collect metrics with Prometheus and visualize them in Grafana. Quake AI does not run a managed observability stack. Deploy the [monitoring stack template](/resources/iac-templates/monitoring-stack) with OpenTofu, install components manually on a VM, or run `kube-prometheus-stack` on Magnum.

<PrerequisiteBlock methods={["terraform", "cli"]}>

- Workloads that expose metrics (node_exporter on VMs, application `/metrics` endpoints, or kube-state-metrics on Kubernetes)
- Network paths from the monitoring hosts to those endpoints (security groups and private subnets)

</PrerequisiteBlock>

## Path A: OpenTofu template (recommended)

1. Clone or copy the [monitoring stack template](/resources/iac-templates/monitoring-stack).
2. Follow the [deploy monitoring stack tutorial](/resources/deployments/deploy-monitoring-stack-template) for variable tuning, apply, Grafana sign-in, and teardown.
3. Add scrape targets for your application instances (private IPs and ports).

The template provisions Prometheus and Grafana on a private subnet with a floating IP for dashboard access.

## Path B: Manual install on a VM

On a dedicated monitoring instance, install Docker first if it is not already present:

```bash
curl -fsSL https://get.docker.com | sudo sh
```

Run node_exporter on each target VM (port 9100):

```bash
sudo docker run -d --net=host --pid=host \
  -v /:/host:ro,rslave prom/node-exporter:latest \
  --path.rootfs=/host
```

Install Prometheus and Grafana with your distribution packages or Docker, then edit `prometheus.yml` `scrape_configs` to list each target's private IP.

## Path C: Kubernetes (kube-prometheus-stack)

On a Magnum cluster with kubeconfig configured ([manage cluster](/docs/kubernetes/how-to/manage-cluster)):

```bash
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm upgrade --install monitoring prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace
```

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the `kube-prometheus-stack` chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin `--version 72.9.1`, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the `--version` flag to install the current chart.

Port-forward Grafana or expose it with a Service and floating IP per your security model.

## Verify

- Prometheus **Targets** UI shows `UP` for each scrape job.
- Grafana dashboards render CPU, memory, and request-rate panels.
- Alertmanager (if enabled) routes test alerts to your on-call channel.

## See also

- [Monitoring overview](/docs/operate/monitoring)
- [Monitoring stack template](/resources/iac-templates/monitoring-stack)
- [Deploy the monitoring stack template with OpenTofu](/resources/deployments/deploy-monitoring-stack-template)
- [How to ship application logs off your VMs](/docs/operate/monitoring/how-to-ship-logs)
