# Monitoring

Source: https://docs.quake.ai/docs/operate/monitoring
Markdown: https://docs.quake.ai/docs/operate/monitoring.md

---

# Monitoring

Monitoring is the practice of collecting, analyzing, and acting on data about your infrastructure and applications. On Quake AI, monitoring is a customer responsibility: the platform provides the compute, network, and storage resources, and you instrument and observe your workloads.

This section covers what to monitor, how to set up collection pipelines, and when to alert.

<DocsSectionLinks section="operate/monitoring" />

## What to monitor

Effective monitoring covers four layers:

- **Infrastructure**: CPU, memory, disk I/O, and network throughput on your instances. These metrics tell you whether your resources are sized correctly and when to scale.
- **Application**: request rates, error rates, latency, and business-specific metrics. These tell you whether your software is working correctly from your users' perspective.
- **Logs**: system logs (`syslog`, `auth.log`), application logs, and audit trails. Logs provide the diagnostic detail that metrics alone cannot capture.
- **Quotas and billing**: resource usage against your project quotas. Unexpected quota consumption may indicate runaway automation or compromised credentials.

## Templates and tutorials

- [Monitoring stack template](/resources/iac-templates/monitoring-stack): deploy Prometheus + Grafana on Quake AI with OpenTofu
- [Deploy the monitoring stack template with OpenTofu](/resources/deployments/deploy-monitoring-stack-template): end-to-end tutorial for the template

## See also

- [Operate overview](/docs/operate)
- [Resource tiers](/docs/account/resource-tiers): project quotas and usage
- [Security hardening checklist](/docs/security/hardening-checklist): access logs and quota auditing
- [Dashboard](/docs/account/dashboard): project-level resource usage in the portal
- [Runbooks](/docs/operate/runbooks)
- [Troubleshooting](/docs/operate/troubleshooting)
