Skip to content

High Availability on Quake AI

Explanation · Updated Jun 2026

High availability on Quake AI

A single instance running your application is a single point of failure. When that instance goes down (planned maintenance, hardware fault, or software crash), your users see an outage. High availability (HA) is the practice of designing systems so that no single failure takes down the service.

HA is an architecture you build from platform primitives. Quake AI gives you the building blocks: anti-affinity server groups, floating IPs, persistent volumes, and isolated networks. You compose them with application-level replication and traffic routing that tolerates failures.

How to think about availability#

Availability is measured as uptime percentage over a period:

TargetAnnual downtimeWhat it means
99% ("two nines")3.65 daysAcceptable for internal tools
99.9% ("three nines")8.76 hoursStandard for most web applications
99.99% ("four nines")52.6 minutesRequires automated failover, no manual steps
99.999% ("five nines")5.26 minutesRequires redundancy at every layer

Each additional nine multiplies the engineering effort. Most production workloads target three or four nines, which is achievable on Quake AI with careful design.

Availability improves when you reduce the blast radius of failures: when a component fails, less of your system is affected, and recovery is faster.

HA patterns on Quake AI#

Anti-affinity server groups#

By default, the scheduler may place multiple instances on the same physical host. If that host fails, the instances on it go down together, defeating the purpose of running replicas.

Anti-affinity server groups instruct the scheduler to place group members on different physical hosts. If one host fails, the other instances in the group remain unaffected.

Use anti-affinity groups for:

  • Application replicas behind a self-managed reverse proxy or external edge service
  • Database primary and replica nodes
  • Kubernetes master nodes (the Kubernetes service uses this automatically for multi-master clusters)

Traffic distribution and failover#

An application designed for high availability needs a traffic entry point that can stop sending requests to an unhealthy replica. Choose a routing layer that matches your architecture:

  • Put a CDN or WAF with health-checked origin pools in front of floating IPs attached to application instances.
  • Run Caddy, nginx, HAProxy, or Traefik on a dedicated reverse-proxy tier. Give each proxy instance a floating IP, and use DNS failover or an external edge service across those addresses.
  • Use DNS health checks and low-TTL records to route clients between floating IPs when your DNS provider supports active failover.

Deploy replicas and reverse-proxy nodes in anti-affinity server groups. Configure application health checks at the routing layer, test the failover path, and keep backend instances on private networks when the proxy tier can reach them.

Volume snapshots for data protection#

Volumes persist independently of instances, but they are still a single copy of your data. For HA, combine volumes with:

  • Regular snapshots: automated point-in-time copies that let you restore if data is corrupted or accidentally deleted
  • Volume clones: create read replicas or standby copies for failover scenarios
  • Application-level replication: for databases, use the database's own replication (PostgreSQL streaming replication, MySQL binary log replication) to maintain a hot standby on a separate instance

Snapshots protect against data loss; replication protects against instance failure. A strong HA design uses both.

Network isolation#

Separate your tiers across dedicated networks to limit the blast radius of failures and contain security incidents:

  • Front-end network: reverse proxies and instances serving public traffic, with floating IPs and tight security groups
  • Application network: application servers communicating with each other and with the data tier, no public access
  • Data network: database instances on an isolated subnet, accessible only from the application tier

Each network has its own security groups, so a compromise or misconfiguration in one tier does not expose the others. Routers connect the tiers where traffic needs to flow.

InternetExternal EdgeFloating IP 1Floating IP 2Front-end NetworkData NetworkApp 1App 2DB PrimaryDB ReplicaVolumeVolume replication
Click to zoom
Multi-tier HA architecture with an external edge, redundant application instances, and network isolation

What the platform handles vs. what you build#

LayerPlatform responsibilityYour responsibility
Physical hardwareRedundant power, network, and storage infrastructureNothing; this is managed
HypervisorHost placement, live migration boundariesUse anti-affinity groups to spread instances
NetworkSoftware-defined networking, router redundancyDesign network topology, security group rules
Traffic routingFloating IPs provide stable addresses for instancesConfigure an external edge, DNS failover, or self-managed reverse proxies and health checks
StorageReplicated block storage backendSnapshots, application-level replication
ApplicationNothing; this is your domainDeploy replicas, handle failover, test on a schedule

Operational considerations#

Test your failover. Kill an instance and verify that traffic reroutes within your acceptable recovery window. Do this on a schedule, not only during initial setup.

Automate recovery. Manual failover procedures add minutes to outages. Use health checks in your edge or reverse-proxy layer, automated snapshot schedules through the CLI or API, and infrastructure-as-code (OpenTofu or Heat) so you can rebuild environments on demand.

Monitor the stack. HA requires visibility. Track instance health, reverse-proxy or edge health, volume I/O latency, and network connectivity. Alert on failures before they become customer reports.

Plan for capacity. HA means running more resources than a single-instance deployment. Budget for at least 2x the compute (active instances plus failover capacity) and additional storage for snapshots and replicas. Check your project quotas before deploying.

Document your architecture. Write down which components are redundant, what the failover path is for each, and what the expected recovery time is. When an outage happens at 3 AM, the runbook is what saves you, not your memory.

Further reading#

On this platform:

  • Server groups: anti-affinity and affinity scheduling policies
  • Volumes: persistent storage and snapshot strategies
  • Networks: private network topology and isolation
  • Floating IPs: stable public endpoints for failover

External resources:

Quick answers

Related content

Was this page helpful?