Runbooks
Overview · Updated May 2026
Runbooks
A runbook is a step-by-step procedure for handling a specific operational scenario. Runbooks reduce mean time to resolution (MTTR) by giving you a tested, repeatable path from symptom to fix; no guesswork under pressure.
Each runbook in this section follows a consistent structure:
- Symptoms: what you observe (error messages, metrics, user reports)
- Diagnosis: how to confirm the root cause
- Resolution: step-by-step fix with commands and expected output
- Verification: how to confirm the issue is resolved
- Prevention: changes to avoid recurrence
- Instance Connectivity Troubleshooting: Diagnose and fix SSH timeouts, refused connections, and network unreachability for Quake AI instances.
- Instance Lifecycle Troubleshooting: Diagnose instances stuck in BUILD, ERROR, or VERIFY_RESIZE states and recover to a working state.
- Kubernetes Cluster Troubleshooting: Diagnose Kubernetes cluster creation failures, stuck pods, networking issues, and storage problems on Quake AI.
- Object Storage Access Troubleshooting: Diagnose S3 403 errors, credential confusion, and upload failures for Quake AI object storage.
- Volume and Storage Troubleshooting: Diagnose volumes stuck in attaching, detaching, or error states and resolve delete conflicts.
Available runbooks#
Additional references#
These existing guides cover supplementary operational procedures:
- Interrupt VM boot process: recover an instance stuck in a boot loop or locked out by a misconfiguration
- Retrieve Windows password: recover access to a Windows Server instance
- Advanced Kubernetes troubleshooting: diagnose control plane and workload issues
- Create an instance snapshot: create a point-in-time backup before maintenance
- Allocate floating IPs: restore public access after a networking change
See also#
See Also
Was this page helpful?