How to monitor your Quake AI workload with Prometheus and Grafana
Coming from another cloud?
▸AWS·Cloudwatch
This Quake AI feature maps to AWS’s Cloudwatch.
▸DigitalOcean·Monitoring
This Quake AI feature maps to DigitalOcean’s Monitoring.
▸Google Cloud·Cloud Monitoring
This Quake AI feature maps to Google Cloud’s Cloud Monitoring.
How to monitor your Quake AI workload with Prometheus and Grafana
Collect metrics with Prometheus and visualize them in Grafana. Quake AI does not run a managed observability stack. Deploy the monitoring stack template with OpenTofu, install components manually on a VM, or run kube-prometheus-stack on Magnum.
Prerequisites
- TerraformOpenTofu installed with Quake AI provider configured
- CLIOpenStack CLI installed and authenticated (
clouds.yamloropenrcsourced)
Windows: CLI examples use bash. Set up a Linux CLI environment on Windows before proceeding.
- Workloads that expose metrics (node_exporter on VMs, application
/metricsendpoints, or kube-state-metrics on Kubernetes) - Network paths from the monitoring hosts to those endpoints (security groups and private subnets)
Path A: OpenTofu template (recommended)#
- Clone or copy the monitoring stack template.
- Follow the deploy monitoring stack tutorial for variable tuning, apply, Grafana sign-in, and teardown.
- Add scrape targets for your application instances (private IPs and ports).
The template provisions Prometheus and Grafana on a private subnet with a floating IP for dashboard access.
Path B: Manual install on a VM#
On a dedicated monitoring instance, install Docker first if it is not already present:
curl -fsSL https://get.docker.com | sudo shRun node_exporter on each target VM (port 9100):
sudo docker run -d --net=host --pid=host \
-v /:/host:ro,rslave prom/node-exporter:latest \
--path.rootfs=/hostInstall Prometheus and Grafana with your distribution packages or Docker, then edit prometheus.yml scrape_configs to list each target's private IP.
Path C: Kubernetes (kube-prometheus-stack)#
On a Magnum cluster with kubeconfig configured (manage cluster):
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm upgrade --install monitoring prometheus-community/kube-prometheus-stack \
--version 72.9.1 \
--namespace monitoring --create-namespaceMatch the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the kube-prometheus-stack chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin --version 72.9.1, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the --version flag to install the current chart.
Port-forward Grafana or expose it with a Service and floating IP per your security model.
Verify#
- Prometheus Targets UI shows
UPfor each scrape job. - Grafana dashboards render CPU, memory, and request-rate panels.
- Alertmanager (if enabled) routes test alerts to your on-call channel.
See also#
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.