Operate
Operate
The Operate section covers monitoring, troubleshooting, maintenance, and operational runbooks for workloads running on Quake AI. This is customer-side operations content: how you keep your infrastructure healthy, diagnose problems, and respond to incidents. For platform-side responsibilities, see the shared responsibility model.
Audience#
This section is for DevOps engineers and administrators managing production workloads on Quake AI. It assumes you are comfortable with Linux system administration, SSH access, and the OpenStack CLI. If you are new to the platform, start with the Quickstart and its learning paths first.
Sections#
Runbooks#
Step-by-step procedures for specific operational scenarios: the kind of document you reach for at 2 AM when something is down. Each runbook covers a single incident type with clear diagnosis steps, resolution actions, and verification.
- Instance connectivity: SSH, floating IP, security groups, DNS, inter-instance
- Instance lifecycle: BUILD stuck, ERROR states, quota failures
- Volume troubleshooting: stuck states, deletion blocks, snapshots, disk full
- Object storage access: auth errors, signing, multipart uploads
- Kubernetes clusters: creation stalls, Heat stacks, kubectl
Monitoring#
Guides for observing your infrastructure: resource utilization, application health, log aggregation, and alerting patterns. Monitoring content covers what to watch, how to set up collection, and when to act.
Troubleshooting#
Systematic approaches to diagnosing common problems: connectivity failures, instance boot issues, storage errors, and Kubernetes cluster problems. Each guide follows a symptom-first structure: what you see, what causes it, how to fix it.
- Quota and limits: diagnose and resolve quota exhaustion
- Auth token diagnostics: 401/403, credential scope, token expiry
- Support ticket evidence: what to collect before escalating
Related resources#
- How to back up and restore a Quake AI VM: snapshot-based VM recovery before maintenance
- How to schedule volume snapshots and restore data: Cinder snapshot schedules for data volumes
- How to monitor your Quake AI workload with Prometheus and Grafana: deploy the monitoring stack template end to end
- How to ship application logs off your VMs: forward application logs to Loki or a SaaS target
- Create an instance snapshot: backup a running VM before maintenance
- Security hardening checklist: audit your project security posture
- Interrupt VM boot process: recover from a misconfigured instance
- Advanced Kubernetes troubleshooting: diagnose cluster and workload issues
- Resource tiers: understand and monitor your quota usage