Monitoring
Overview · Updated May 2026
Coming from another cloud?
▸AWS·Cloudwatch
This Quake AI feature maps to AWS’s Cloudwatch.
▸DigitalOcean·Monitoring
This Quake AI feature maps to DigitalOcean’s Monitoring.
Monitoring
Monitoring is the practice of collecting, analyzing, and acting on data about your infrastructure and applications. On Quake AI, monitoring is a customer responsibility: the platform provides the compute, network, and storage resources, and you instrument and observe your workloads.
This section covers what to monitor, how to set up collection pipelines, and when to alert.
- How to Monitor Your Quake AI Workload with Prometheus and Grafana: Collect metrics with Prometheus and visualize them in Grafana. Quake AI does not run a managed observability stack. Deploy the monitoring stack template with OpenTofu, install components manually on a...
- How to Ship Application Logs Off Your VMs: Forward application and system logs to a central store for search and alerting. Quake AI does not provide a managed log service. Run self-hosted Loki on a Quake AI VM, or forward to a SaaS log...
What to monitor#
Effective monitoring covers four layers:
- Infrastructure: CPU, memory, disk I/O, and network throughput on your instances. These metrics tell you whether your resources are sized correctly and when to scale.
- Application: request rates, error rates, latency, and business-specific metrics. These tell you whether your software is working correctly from your users' perspective.
- Logs: system logs (
syslog,auth.log), application logs, and audit trails. Logs provide the diagnostic detail that metrics alone cannot capture. - Quotas and billing: resource usage against your project quotas. Unexpected quota consumption may indicate runaway automation or compromised credentials.
Templates and tutorials#
- Monitoring stack template: deploy Prometheus + Grafana on Quake AI with OpenTofu
- Deploy the monitoring stack template with OpenTofu: end-to-end tutorial for the template
See also#
- Operate overview
- Resource tiers: project quotas and usage
- Security hardening checklist: access logs and quota auditing
- Dashboard: project-level resource usage in the portal
- Runbooks
- Troubleshooting
See Also
Was this page helpful?