Skip to content

How to monitor your Quake AI workload with Prometheus and Grafana

How-to

Coming from another cloud?

▸AWS·Cloudwatch

This Quake AI feature maps to AWS’s Cloudwatch.

▸DigitalOcean·Monitoring

This Quake AI feature maps to DigitalOcean’s Monitoring.

▸Google Cloud·Cloud Monitoring

This Quake AI feature maps to Google Cloud’s Cloud Monitoring.

How to monitor your Quake AI workload with Prometheus and Grafana

Collect metrics with Prometheus and visualize them in Grafana. Quake AI does not run a managed observability stack. Deploy the monitoring stack template with OpenTofu, install components manually on a VM, or run kube-prometheus-stack on Magnum.

Prerequisites

Windows: CLI examples use bash. Set up a Linux CLI environment on Windows before proceeding.

  • Workloads that expose metrics (node_exporter on VMs, application /metrics endpoints, or kube-state-metrics on Kubernetes)
  • Network paths from the monitoring hosts to those endpoints (security groups and private subnets)
  1. Clone or copy the monitoring stack template.
  2. Follow the deploy monitoring stack tutorial for variable tuning, apply, Grafana sign-in, and teardown.
  3. Add scrape targets for your application instances (private IPs and ports).

The template provisions Prometheus and Grafana on a private subnet with a floating IP for dashboard access.

Path B: Manual install on a VM#

On a dedicated monitoring instance, install Docker first if it is not already present:

bash
curl -fsSL https://get.docker.com | sudo sh

Run node_exporter on each target VM (port 9100):

bash
sudo docker run -d --net=host --pid=host \
  -v /:/host:ro,rslave prom/node-exporter:latest \
  --path.rootfs=/host

Install Prometheus and Grafana with your distribution packages or Docker, then edit prometheus.yml scrape_configs to list each target's private IP.

Path C: Kubernetes (kube-prometheus-stack)#

On a Magnum cluster with kubeconfig configured (manage cluster):

bash
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm repo update
helm upgrade --install monitoring prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the kube-prometheus-stack chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin --version 72.9.1, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the --version flag to install the current chart.

Port-forward Grafana or expose it with a Service and floating IP per your security model.

Verify#

  • Prometheus Targets UI shows UP for each scrape job.
  • Grafana dashboards render CPU, memory, and request-rate panels.
  • Alertmanager (if enabled) routes test alerts to your on-call channel.

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

For the full policy, see Usage Guidelines.

Was this page helpful?