A single instance running your application is a single point of failure. When that instance goes down (planned maintenance, hardware fault, or software crash), your users see an outage. High availability (HA) is the practice of designing systems so that no single failure takes down the service.
HA is an architecture you build from platform primitives. Quake AI gives you the building blocks: anti-affinity server groups, floating IPs, persistent volumes, and isolated networks. You compose them with application-level replication and traffic routing that tolerates failures.
Availability is measured as uptime percentage over a period:
Target
Annual downtime
What it means
99% ("two nines")
3.65 days
Acceptable for internal tools
99.9% ("three nines")
8.76 hours
Standard for most web applications
99.99% ("four nines")
52.6 minutes
Requires automated failover, no manual steps
99.999% ("five nines")
5.26 minutes
Requires redundancy at every layer
Each additional nine multiplies the engineering effort. Most production workloads target three or four nines, which is achievable on Quake AI with careful design.
Availability improves when you reduce the blast radius of failures: when a component fails, less of your system is affected, and recovery is faster.
By default, the scheduler may place multiple instances on the same physical host. If that host fails, the instances on it go down together, defeating the purpose of running replicas.
Anti-affinity server groups instruct the scheduler to place group members on different physical hosts. If one host fails, the other instances in the group remain unaffected.
Use anti-affinity groups for:
Application replicas behind a self-managed reverse proxy or external edge service
Database primary and replica nodes
Kubernetes master nodes (the Kubernetes service uses this automatically for multi-master clusters)
An application designed for high availability needs a traffic entry point that can stop sending requests to an unhealthy replica. Choose a routing layer that matches your architecture:
Put a CDN or WAF with health-checked origin pools in front of floating IPs attached to application instances.
Run Caddy, nginx, HAProxy, or Traefik on a dedicated reverse-proxy tier. Give each proxy instance a floating IP, and use DNS failover or an external edge service across those addresses.
Use DNS health checks and low-TTL records to route clients between floating IPs when your DNS provider supports active failover.
Deploy replicas and reverse-proxy nodes in anti-affinity server groups. Configure application health checks at the routing layer, test the failover path, and keep backend instances on private networks when the proxy tier can reach them.
Volumes persist independently of instances, but they are still a single copy of your data. For HA, combine volumes with:
Regular snapshots: automated point-in-time copies that let you restore if data is corrupted or accidentally deleted
Volume clones: create read replicas or standby copies for failover scenarios
Application-level replication: for databases, use the database's own replication (PostgreSQL streaming replication, MySQL binary log replication) to maintain a hot standby on a separate instance
Snapshots protect against data loss; replication protects against instance failure. A strong HA design uses both.
Separate your tiers across dedicated networks to limit the blast radius of failures and contain security incidents:
Front-end network: reverse proxies and instances serving public traffic, with floating IPs and tight security groups
Application network: application servers communicating with each other and with the data tier, no public access
Data network: database instances on an isolated subnet, accessible only from the application tier
Each network has its own security groups, so a compromise or misconfiguration in one tier does not expose the others. Routers connect the tiers where traffic needs to flow.
Click to zoom
Multi-tier HA architecture with an external edge, redundant application instances, and network isolation
Test your failover. Kill an instance and verify that traffic reroutes within your acceptable recovery window. Do this on a schedule, not only during initial setup.
Automate recovery. Manual failover procedures add minutes to outages. Use health checks in your edge or reverse-proxy layer, automated snapshot schedules through the CLI or API, and infrastructure-as-code (OpenTofu or Heat) so you can rebuild environments on demand.
Monitor the stack. HA requires visibility. Track instance health, reverse-proxy or edge health, volume I/O latency, and network connectivity. Alert on failures before they become customer reports.
Plan for capacity. HA means running more resources than a single-instance deployment. Budget for at least 2x the compute (active instances plus failover capacity) and additional storage for snapshots and replicas. Check your project quotas before deploying.
Document your architecture. Write down which components are redundant, what the failover path is for each, and what the expected recovery time is. When an outage happens at 3 AM, the runbook is what saves you, not your memory.