Skip to content

Operate

Overview · Updated May 2026

Operate

The Operate section covers monitoring, troubleshooting, maintenance, and operational runbooks for workloads running on Quake AI. This is customer-side operations content: how you keep your infrastructure healthy, diagnose problems, and respond to incidents. For platform-side responsibilities, see the shared responsibility model.

Audience#

This section is for DevOps engineers and administrators managing production workloads on Quake AI. It assumes you are comfortable with Linux system administration, SSH access, and the OpenStack CLI. If you are new to the platform, start with the Quickstart and its learning paths first.

Sections#

Runbooks#

Step-by-step procedures for specific operational scenarios: the kind of document you reach for at 2 AM when something is down. Each runbook covers a single incident type with clear diagnosis steps, resolution actions, and verification.

Monitoring#

Guides for observing your infrastructure: resource utilization, application health, log aggregation, and alerting patterns. Monitoring content covers what to watch, how to set up collection, and when to act.

Troubleshooting#

Systematic approaches to diagnosing common problems: connectivity failures, instance boot issues, storage errors, and Kubernetes cluster problems. Each guide follows a symptom-first structure: what you see, what causes it, how to fix it.

Was this page helpful?