Understanding limitations and advanced features of Quake AI Kubernetes clusters
Coming from another cloud?
▸AWS·Amazon EKS Cluster
Amazon EKS Cluster
- EKS control plane fully AWS-managed, single-tenant, across 3 AZs with auto scale/replace; on Quake AI you run the control plane on Nova instances, provisioned through Magnum or self-managed with OpenTofu, and you operate it.
- EKS regional API endpoint with SLA; Quake AI exposes the kube API via Neutron LB with floating IP.
- EKS charges a per-hour cluster platform fee on top of the underlying compute; Quake AI charges only for underlying Nova/Neutron/Cinder resources with no K8s platform fee.
- EKS managed nodes auto AMI updates, Spot integration; Quake AI self-managed nodes require manual OS image selection and update management.
▸Azure·AKS Cluster
AKS Cluster
- Azure automatically provisions and manages the control plane at no additional cost (Free tier) or fixed fee (Standard tier with SLA), offloading health monitoring and upgrades; on Quake AI you provision a cluster through Magnum (openstack coe cluster create) or self-managed Kubernetes on Nova instances (OpenTofu plus kubeadm, k3s, or RKE2), and you operate the cluster after creation.
- No OpenStack integration; uses Azure Resource Manager for cluster lifecycle.
- Pre-configured with Azure-specific defaults and add-ons like application routing.
- Managed via Azure Virtual Machine Scale Sets (VMSS) with auto-scaling and upgrades; Quake AI uses Nova instances provisioned via OpenTofu with user-managed scaling and upgrades.
▸DigitalOcean·Kubernetes Cluster
Kubernetes Cluster
- Uses proprietary DigitalOcean API (POST /v2/kubernetes/clusters) for provisioning; Quake AI users provision Nova instances via OpenTofu and bootstrap Kubernetes manually (kubeadm, k3s).
- No cluster templates; direct specification of node pools in create request. Quake AI uses OpenTofu modules or IaC patterns to define cluster topology.
- Fully managed control plane at no cost (HA mode is a paid optional add-on); on Quake AI you provision the control plane through Magnum or self-managed on Nova instances and operate it yourself, including HA configuration.
- Node pools consist of identical Droplet-based worker nodes managed solely via kubectl (no SSH access); Quake AI nodes are Nova instances with full SSH access via keypair.
▸Google Cloud·GKE Cluster
GKE Cluster
- GKE provides Autopilot mode with fully managed node provisioning and scaling by Google; on Quake AI you provision clusters through Magnum or self-managed Kubernetes on Nova instances and manage node scaling yourself.
- Control plane is fully managed with automatic upgrades through release channels; Quake AI requires user-provisioned and user-managed control plane nodes.
- Cluster creation uses gcloud CLI vs OpenStack CLI (openstack coe cluster create).
- Custom machine types, spot VMs, accelerators in node pools; Quake AI K8s nodes use standard Nova flavors.
▸Hetzner·No managed Kubernetes service (self-managed with CCM and CSI)
No managed Kubernetes service (self-managed with CCM and CSI)
- Similar self-managed model to Quake AI; both require users to provision servers and install Kubernetes with tools like OpenTofu/Terraform and k3s/kubeadm. Hetzner provides a Cloud Controller Manager and CSI driver, while Quake AI uses OpenStack-compatible storage and network integrations.
- Hetzner-specific cloud-provider=external CCM must be deployed manually in kube-system namespace.
- No cluster API or web UI for cluster lifecycle; all via Hetzner Cloud API token and kubectl.
Understanding limitations and advanced features of Quake AI Kubernetes clusters
Floating IPs and Kubernetes LoadBalancer services both translate addresses. When you assign a floating IP to a cluster node and also expose a LoadBalancer service, the two NAT paths can conflict. You then see failed health checks and unreachable services.
The problem: IP conflict between floating IPs and load balancers#
Assigning floating IPs directly to cluster nodes while you also deploy LoadBalancer-type services creates overlapping address translation:
- A floating IP translates public traffic to the node's internal IP.
- A load balancer performs NAT toward service endpoints (pods or nodes).
When both run at once, NAT rules can collide and the load balancer stops passing traffic.
Recommended workarounds#
Avoid assigning floating IPs to nodes at create time#
Quake AI Magnum templates leave per-node floating IPs off by default (floating_ip_enabled: false). Magnum still exposes the create-time switches:
- CLI:
--floating-ip-enabledand--floating-ip-disabledonopenstack coe cluster create - API:
floating_ip_enabled
Do not pass --floating-ip-enabled, and do not set floating_ip_enabled to true. Use --floating-ip-disabled, or leave floating_ip_enabled false, so Magnum does not assign a floating IP to each node during installation.
A template may still set master_lb_floating_ip_enabled: true so the cluster API load balancer has a public address. That setting is separate from per-node floating IPs.
Remove floating IPs from nodes on existing clusters#
If an existing cluster already has floating IPs on Kubernetes nodes, disassociate those floating IPs from the nodes. Use the Quake AI web interface or the OpenStack CLI.
Disassociating node floating IPs also removes:
- Direct SSH to nodes over a public IP
- NodePort access to Kubernetes services from the public internet
Avoid re-associating floating IPs#
Temporarily re-associating and then disassociating floating IPs can look like it clears the conflict. The conflict can return after you re-associate them.
Exclude specific nodes from load balancers#
If some nodes must keep a floating IP, exclude those nodes from external load balancers. Apply the node.kubernetes.io/exclude-from-external-load-balancers=true label. See the Kubernetes labels reference for node.kubernetes.io/exclude-from-external-load-balancers for semantics and when to use the label.
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.
Last validated: 01.09.2026