# Kubernetes FAQ

Source: https://docs.quake.ai/docs/kubernetes/faq
Markdown: https://docs.quake.ai/docs/kubernetes/faq.md
> Frequently asked questions about Quake AI Kubernetes: cluster provisioning via Magnum, cluster templates, lifecycle, high availability, migration from EKS/AKS/GKE/DOKS/Hetzner, and troubleshooting.

---

{/*
================================================================================
FAQ: Kubernetes service
================================================================================
Confidence tiers:

  T1  High-confidence: direct restatement of canonical content.
  T2  Medium-confidence: synthesized across multiple pages or implied.
  T3  Validation-needed: pricing, quotas, SLA, support, needs team sign-off.

Every answer ends with a "Source:" line citing the canonical doc(s).
================================================================================
*/}

# Kubernetes FAQ

Frequently asked questions about Kubernetes clusters on Quake AI. For step-by-step instructions see the [Kubernetes how-to guides](/docs/kubernetes). For deeper background see the [concept pages](/docs/kubernetes/concepts/kubernetes). For migration from another provider see the [migration guides](/docs/kubernetes/migration).

## Getting started





### Q: What is the Kubernetes service on Quake AI? [T1]


The Kubernetes service automates cluster provisioning on top of existing Compute, Network, and Storage infrastructure. You define the cluster shape through a reusable **cluster template**, and the platform provisions master VMs, worker VMs, a dedicated cluster network, stable endpoints for the Kubernetes API and etcd, and the router connecting the cluster to the external network. Provisioning is backed by [OpenStack Magnum](https://docs.openstack.org/magnum/latest/) and orchestrated through the Automation service (Heat).

Once the cluster is running, you manage workloads with standard tools: `kubectl`, Helm, and the OpenStack CLI for cluster-level operations such as scaling or deletion. The platform handles the infrastructure setup; you operate the cluster after provisioning, including control plane availability and upgrades.

---





### Q: Is this a managed Kubernetes control plane? [T1]


Quake AI Kubernetes provisions the control plane inside **your project**. The master nodes run as Compute instances, so your team operates control plane availability, etcd backups, and certificate rotation. The platform automates the provisioning of those master VMs and the networking between them.

This model is for teams that want Kubernetes infrastructure provisioning automated without setting up `kubeadm`, CNI, control-plane endpoints, and storage drivers by hand.

---





### Q: What infrastructure does a cluster provision inside my project? [T1]


When you create a cluster, the Kubernetes service provisions the following resources in your project:

| Resource | Role |
|---|---|
| **Master VMs** | Run the Kubernetes API server, scheduler, controller manager, and etcd. Provisioned as Compute instances. |
| **Worker VMs** | Run the kubelet and container runtime. Your application pods are scheduled here. |
| **Network** | A dedicated network and subnet for intra-cluster communication. |
| **Control-plane endpoints** | A stable Kubernetes API endpoint when master load balancing is enabled. Clusters with multiple master nodes also use a stable etcd endpoint. |
| **Router** | Connects the cluster network to the external network. |
| **Persistent volumes** | Backed by Block Storage volumes through the Cinder CSI driver. |

Underlying VMs use Fedora CoreOS images with the container runtime and Kubernetes components pre-installed. The Heat stack that drives provisioning is visible under **Automation** > **Stacks** in the console.

---





### Q: What are the minimum resource requirements to create a cluster? [T1]


Kubernetes clusters consume Compute, Network, and Storage quotas. Minimum requirements by cluster size:

| Cluster size | Master nodes | Worker nodes | Min vCPU | Min RAM | Min block storage | Floating IPs |
|---|---|---|---|---|---|---|
| Minimal (dev/test) | 1 | 1 | 4 shared | 4 GB | 180 GB | 3 |
| Small production | 3 | 2 | 34 dedicated | 80 GB | 500 GB | 4 |
| Medium production | 3 | 5 | 46 dedicated | 128 GB | 1 TB | 7 |

Check your quota before creating a cluster with `openstack quota show` or navigate to **Home** > **Overview** in the console.



A single-master cluster cannot be upgraded to multiple masters later. If you plan to run production workloads, start with three master nodes.



---





## Core concepts





### Q: What is a cluster template, and why do I need one? [T1]


A **cluster template** is a reusable blueprint that captures every parameter the Kubernetes service needs to provision a cluster:

- **Kubernetes version** and container orchestration engine (COE)
- **Compute flavors** for master and worker nodes
- **Network driver** (Flannel or Calico) and external network
- **Volume driver** for persistent storage (Cinder)
- **TLS settings**, Docker storage driver, and registry configuration
- **Labels** that control feature flags (autoscaler, dashboard, network policies)

You create a template once and reuse it for multiple clusters. Platform-provided templates (for example, `Standard-V2.0-k8s-calico-fc38_v1.24.16`) are available by default. Separating the template from the cluster lets you standardize configurations across teams and environments, for instance, a production template with three master nodes and Calico, and a development template with a single master and Flannel.

---





### Q: Which network driver should I choose: Calico or Flannel? [T1]


Choose based on your need for Kubernetes `NetworkPolicy` enforcement:

- **Calico**: recommended for production. Supports Kubernetes `NetworkPolicy` for pod-to-pod traffic control. Configured in VXLAN mode on Quake AI migrations.
- **Flannel**: simpler overlay network with lower operational overhead. Does not enforce `NetworkPolicy`. Suitable for development clusters or workloads that do not require pod-level traffic isolation.

The network driver is set in the cluster template and cannot be changed after the cluster is created.

---





### Q: What does the cluster lifecycle look like? [T1]


| Phase | Status | What happens |
|---|---|---|
| **Create** | `CREATE_IN_PROGRESS` → `CREATE_COMPLETE` | The service provisions VMs, networking, and control-plane endpoints from the cluster template. Creation typically takes 5 to 15 minutes. |
| **Run** | `ACTIVE` / health status | The cluster is active. Interact with it via `kubectl` using a kubeconfig retrieved from the CLI. |
| **Scale** | `UPDATE_IN_PROGRESS` → `UPDATE_COMPLETE` | Add or remove worker nodes. The platform provisions or removes VMs accordingly. |
| **Update** | `UPDATE_IN_PROGRESS` → `UPDATE_COMPLETE` | Modify cluster properties such as node count or labels. |
| **Delete** | `DELETE_IN_PROGRESS` → `DELETE_COMPLETE` | The platform tears down the associated VMs, networks, control-plane endpoints, and Heat stack. |

Cluster health is monitored continuously. The console displays both the operational status and a health status based on node and control plane checks.

---





### Q: How do I achieve high availability for the Kubernetes control plane? [T1]


Create the cluster with **3 master nodes**. Quake AI places master VMs on separate physical hosts using Nova anti-affinity scheduling (the same server-group mechanism available for Compute instances). The cluster uses stable, load-balanced endpoints for the Kubernetes API and etcd across the three masters.



Master nodes cannot be added after cluster creation. If you want a fault-tolerant control plane (3 masters, survives one master failure), set `--master-count 3` at creation time. A single-master cluster cannot be upgraded to multi-master.



For information on how anti-affinity server groups work at the Compute level, see [Server groups](/docs/compute/concepts/server-groups).

---





### Q: What Kubernetes objects and tools are fully portable when migrating? [T1]


The following resources migrate without modification: Deployments, StatefulSets, DaemonSets, ConfigMaps, Secrets (values), Namespaces, ServiceAccounts (minus cloud IAM annotations), all RBAC resources, `NetworkPolicy`, HPA, VPA, `PodDisruptionBudget`, CronJobs, and Jobs.

Helm charts using standard Kubernetes resources deploy cleanly with updated values. The following operators and Helm charts are also fully portable: cert-manager, external-dns, kube-prometheus-stack, Loki, Grafana, Velero, Istio, Linkerd, ArgoCD, Flux, ingress-nginx, Traefik, Metrics Server, and HashiCorp Vault.

Items that need adaptation: Ingress class and annotations, `Service type: LoadBalancer` annotations (remove provider-specific ones), and `PersistentVolumeClaim` StorageClass references (change to `cinder-flash`).

---





## Operations and lifecycle





### Q: How do I create a cluster? [T1]


**Prerequisites:** a cluster template, an SSH key pair, and sufficient project quota.

**Console:** **Services** > **Kubernetes** > **Clusters** > **Create Cluster**. Select a template, set master count (use 3 for production), set worker count, configure networking, and optionally enable Auto Healing and Auto Scaling. Creation takes 5 to 15 minutes.

**CLI:**

```bash
openstack coe cluster create MY_CLUSTER_NAME \
  --cluster-template MY_TEMPLATE_NAME \
  --keypair MY_KEYPAIR \
  --master-count 3 \
  --node-count 2 \
  --timeout 60
```

Monitor status with `openstack coe cluster show MY_CLUSTER_NAME -f value -c status`. Wait for `CREATE_COMPLETE`.

---






Magnum cluster creation on Quake AI requires two adjustments the upstream defaults do not provide: a `boot_volume_size` label on the cluster template, and flavor choices that fit the project's `compute_units` quota.

**Boot volume size.** Every Quake AI flavor has `disk: 0`. Magnum's upstream default produces image-backed (zero-disk) servers, which the platform rejects with `Forbidden: Only volume-backed servers are allowed for flavors with zero disk`. Add `boot_volume_size=40` (the platform-template default) to the cluster template's `--labels` argument:

```bash
openstack coe cluster template create MY_TEMPLATE_NAME \
  --coe kubernetes \
  --image FedoraCoreOS-38 \
  --keypair MY_KEYPAIR \
  --flavor c2a.large \
  --master-flavor c2a.xlarge \
  --external-network PublicStatic \
  --network-driver calico \
  --volume-driver cinder \
  --master-lb-enabled \
  --labels boot_volume_size=40,master_lb_floating_ip_enabled=true
```

**Compute units quota.** The platform's `compute_units` quota caps how many Compute units a project can provision at once. Default-tier projects start at `8000`. Each flavor carries a per-instance reservation in its `rumble:compute_units` property:

| Flavor | vCPUs | RAM | `compute_units` |
|---|---|---|---|
| `c2a.large` | 2 | 4 GB | 2000 |
| `c2a.xlarge` | 4 | 8 GB | 4000 |
| `c2a.2xlarge` | 8 | 16 GB | 8000 |

A cluster with a single `c2a.xlarge` master and a single `c2a.large` worker consumes 6000 `compute_units`, fitting the default quota. A `c2a.2xlarge` master alone consumes 8000 `compute_units` and leaves no headroom for any other resource in the project. Verify your quota with `openstack quota show -f value -c compute_units` before sizing, and request an increase through the Compute Service Request flow before going larger.

The platform-provided cluster templates (for example, `Standard-V2.0-k8s-calico-fc38_v1.24.16`) already carry the four required labels and a working flavor sizing. Cloning a platform template is the fastest path to a working cluster.

---





### Q: Why does `openstack coe cluster create` fail with a Keystone trust or unauthorized error when I use an application credential? [T2]


Magnum cluster creation requires a password-scoped session. `openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Keystone application credentials cannot delegate trust, so the call fails before any Magnum work begins. The cluster transitions to `CREATE_FAILED` with this error:

```text
Resource CREATE failed: Forbidden: resources.kube_cluster_deploy:
Failed to create trustee or trust for Cluster
```

Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` (password auth) before running `openstack coe cluster create`:

```bash
export OS_AUTH_URL=https://keystone.us-east-1.rumble.cloud/v3
export OS_PROJECT_ID=YOUR_PROJECT_ID
export OS_USERNAME=YOUR_USERNAME
export OS_PASSWORD=YOUR_PASSWORD
export OS_USER_DOMAIN_NAME=Default
export OS_PROJECT_DOMAIN_NAME=Default
export OS_IDENTITY_API_VERSION=3
openstack coe cluster create ...
```

The Console wizard at **Services** > **Kubernetes** > **Clusters** > **Create Cluster** works from any logged-in user session because it uses the user's password-scoped token. The same constraint applies to Terraform and Heat templates that drive Magnum: configure the OpenStack provider with password auth before applying a template that creates a cluster.

Application credentials remain the right choice for other workloads on Quake AI, including Compute, Network, Volume, Object Storage, and DNS. The Magnum trust-delegation requirement is the documented exception.

---





### Q: How do I scale worker nodes? [T1]


Scaling only applies to **worker nodes**. Master node count is fixed at creation time.

**Console:** **Services** > **Kubernetes** > **Clusters** > **Resize Cluster** > enter the new node count > **OK**. The cluster status moves to `UPDATE_IN_PROGRESS` while VMs are provisioned or removed.

**CLI:**

```bash
openstack coe cluster resize MY_CLUSTER_NAME NODE_COUNT
```

Monitor with `openstack coe cluster show MY_CLUSTER_NAME -f value -c status` until status returns to `UPDATE_COMPLETE`.

---





### Q: How do I upgrade Kubernetes versions? [T1]


The Kubernetes service does not support in-place version upgrades. The recommended approach is:

1. Create a new cluster template that specifies the target Kubernetes version.
2. Provision a new cluster from that template.
3. Migrate workloads to the new cluster using Velero or by redeploying manifests.
4. Delete the old cluster once traffic is shifted and data is confirmed intact.

This approach keeps the old cluster available for rollback during the migration window.

---





### Q: How do I access my cluster with kubectl? [T1]


Use the CLI to retrieve the kubeconfig file after the cluster reaches `CREATE_COMPLETE`:

```bash
openstack coe cluster config MY_CLUSTER_NAME
```

This outputs an `export KUBECONFIG=...` command. Run it, then verify connectivity:

```bash
kubectl get nodes
```

All master and worker nodes should show `Ready` status. The console does not provide direct `kubectl` access; use the CLI method.

---





### Q: How do I delete a cluster? [T1]


Deleting a cluster removes its VMs, networks, control-plane endpoints, routers, and underlying Heat stack.



Cluster deletion is irreversible. Back up data stored on persistent volumes before deleting. Volumes created through Kubernetes PersistentVolumeClaims with `reclaimPolicy: Delete` are removed with the cluster.



**Console:** **Services** > **Kubernetes** > **Clusters** > **Delete** > confirm.

**CLI:**

```bash
openstack coe cluster delete MY_CLUSTER_NAME
```

If deletion fails with `DELETE_FAILED`, check `openstack coe cluster show MY_CLUSTER_NAME -f value -c status_reason`. Common causes include orphaned resources or dependency conflicts.

---





## Troubleshooting





### Q: Why is my cluster stuck in CREATE_IN_PROGRESS? [T1]


A cluster that stays in `CREATE_IN_PROGRESS` for longer than 30 minutes typically indicates one of four root causes:

1. **Heat stack failure**: inspect the underlying stack with `openstack coe cluster show YOUR_CLUSTER -f value -c stack_id`, then `openstack stack show YOUR_STACK_ID`. If the stack is `CREATE_FAILED`, read `resource_status_reason`.
2. **WaitCondition timeout**: a node-level service (etcd, kubelet, or the Kubernetes API server) did not signal readiness in time. Check the instance console log for cloud-init errors, metadata timeouts, or image pull failures.
3. **Quota or capacity limits**: clusters consume multiple instances, volumes, ports, and often floating IPs. Verify project quotas with `openstack quota show` and check availability zone capacity.
4. **Network or DNS blocking bootstrap**: nodes must reach the metadata service and any image registries referenced by the cluster template. Confirm DNS nameservers are set correctly on the cluster subnet.

If the Heat stack is `CREATE_COMPLETE` but Magnum still shows `CREATE_IN_PROGRESS`, wait a few minutes for Magnum to reconcile. If it does not update, treat the cluster as stuck and delete it.

---





### Q: Why can't kubectl connect to my cluster after CREATE_COMPLETE? [T1]


`kubectl get nodes` returning "Unable to connect to the server" or a timeout despite the cluster showing `CREATE_COMPLETE` has four common causes:

1. **Stale or hand-edited kubeconfig**: regenerate with `openstack coe cluster config YOUR_CLUSTER --dir ~/`. This is the most common cause.
2. **API server not on a reachable IP**: open `~/config` and verify the `server:` URL uses a public floating IP, not an unreachable private address.
3. **Security group blocking TCP 6443**: confirm the cluster's API security group allows ingress on TCP port 6443 from your IP or trusted CIDR.
4. **Master instances down**: run `openstack server list | grep YOUR_CLUSTER` to confirm masters are `ACTIVE`, not `SHUTOFF` or in `ERROR`.

Test raw API reachability with `curl -k https://API_SERVER_IP:6443/version`. A JSON response confirms the API is up; a timeout points to routing, floating IP, or security group issues.

---





### Q: Why does my GitHub Actions or GitLab CI job fail to run `openstack coe` or `kubectl` on a Magnum cluster? [T2]


CI images often ship `python-openstackclient` without the Magnum (`coe`) plugin, or a slim Python base without `kubectl` and Helm. Install the full client stack in the job before you call cluster APIs:

```bash
pip install python-openstackclient python-magnumclient
```

For GitLab `python:3.12-slim` and similar images, also install `kubectl` and Helm from their upstream releases (or use a runner image that already includes them). Authenticate with a password-scoped OpenStack session when the workflow creates or reconfigures clusters; application credentials cannot delegate the Keystone trust Magnum requires.





` for several minutes?"
  tier="T2"
  service="kubernetes"
  concept={["floating_ips", "kubernetes"]}
  error_pattern="k8s-loadbalancer-external-ip-pending"
  affected_methods={["cli"]}
  sources={[
    { label: "Deploy a Helm chart", href: "/docs/kubernetes/how-to/deploy-helm-chart" },
    { label: "Kubernetes cluster troubleshooting", href: "/docs/operate/runbooks/kubernetes-troubleshooting" },
  ]}
>
A `Service` with `type: LoadBalancer` stays `<pending>` while the cluster provisions its public endpoint and floating IP. Provisioning usually completes within a few minutes; long delays often trace to floating IP quota limits, a subnet without a router to `PublicStatic`, or a cluster template with `floating_ip_enabled: false`.

Run `kubectl describe svc SERVICE_NAME -n NAMESPACE` and check Events for cloud-controller or networking errors. Confirm the project still has floating IP capacity with `openstack quota show --usage`. Regenerate kubeconfig with `openstack coe cluster config CLUSTER_NAME --dir "$HOME/.kube/CLUSTER_NAME"` so Helm and kubectl target the cluster you tested.





## Migration





### Q: What is the recommended path for migrating Kubernetes workloads to Quake AI? [T1]


The migration guides all share a common pattern:

1. **Audit** your source cluster: export all manifests, identify provider-specific resources (cloud CSI drivers, cloud controller manager, ingress controllers, pod identity mechanisms).
2. **Provision** a new self-managed Kubernetes cluster on Nova instances using OpenTofu with the `openstack` provider, then install RKE2 (recommended for production) or k3s.
3. **Install the OpenStack cloud-provider stack:** `openstack-cloud-controller-manager`, Cinder CSI driver with `cinder-flash` StorageClass, and Calico (VXLAN) or Cilium as the CNI.
4. **Adapt** provider-specific resources: replace the cloud CSI driver with Cinder CSI, replace the cloud ingress controller with nginx-ingress, replace pod-identity mechanisms with OpenStack application credentials in Kubernetes Secrets.
5. **Migrate data** with [Velero](https://velero.io/) using `--default-volumes-to-fs-backup` and StorageClass remapping to `cinder-flash`.
6. **Validate** workloads and cut over DNS.



Magnum is the other supported provisioning path: it builds a cluster from a curated template instead of installing Kubernetes yourself. The steps above use the self-managed path because it lets you match a source cluster's exact Kubernetes version and CNI; choose Magnum when the template catalog covers your requirements.



---





### Q: What is the recommended migration path from AWS EKS? [T1]


**Complexity: Medium-High. Estimated timeline: 4 to 8 weeks.**

Key tasks unique to EKS:

- **IRSA removal**: remove `eks.amazonaws.com/role-arn` annotations from every ServiceAccount; replace with OpenStack [application credentials](/docs/identity/how-to/create-application-credential) in Kubernetes Secrets.
- **EBS to Cinder**: change `storageClassName` from `gp2`/`gp3` to `cinder-flash` in all PVCs; migrate data with Velero filesystem backup.
- **ALB to nginx-ingress**: replace ALB annotations with nginx-ingress annotations; TLS termination moves from ACM to cert-manager + Let's Encrypt.
- **VPC CNI replacement**: replace `aws-node` DaemonSet with Calico or Cilium.
- **Fargate pods**: remove `eks.amazonaws.com/compute-type: fargate` annotations; deploy to regular Nova-backed node pools.
- **ECR images**: mirror all ECR images to Docker Hub or Harbor before cutover; ECR auth tokens expire every 12 hours.
- **Karpenter**: remove all Karpenter CRDs and NodePools; replace with Cluster Autoscaler with `--cloud-provider=openstack` or manual Nova scaling.

See also the [Coming from AWS: Kubernetes section](/resources/migration/coming-from-aws) for a service-by-service terminology mapping.

---





### Q: What is the recommended migration path from Azure AKS? [T1]


**Complexity: High. Estimated timeline: 6 to 12 weeks.** This is the most complex of the five migrations.

Key tasks unique to AKS:

- **Export Azure Key Vault secrets** first, pod startup fails without Azure network access, so copy all secrets to Kubernetes Secrets (or HashiCorp Vault) before migrating pods.
- **Entra ID Workload Identity / AAD Pod Identity removal**: remove `azure.workload.identity/client-id` annotations; remove `AzureIdentity` and `AzureIdentityBinding` CRDs and the NMI DaemonSet.
- **AGIC to nginx-ingress**: replace Application Gateway Ingress Controller annotations.
- **Azure CNI to Calico/Cilium**: Azure CNI uses VNET secondary IPs; the new pod CIDR changes, so update all firewall rules.
- **Azure Managed Disk CSI to Cinder**: change `storageClassName` from `managed-csi`/`managed-csi-premium` to `cinder-flash`.
- **Azure Monitor DaemonSets**: remove `omsagent`/`ama-logs` DaemonSets **before** deploying Prometheus to avoid metrics duplication.
- **Azure DevOps pipelines**: create a kubeconfig-based Kubernetes service connection pointing to the Quake AI cluster.

---





### Q: What is the recommended migration path from GCP GKE? [T1]


**Complexity: Medium-High. Estimated timeline: 3 to 6 weeks** (longer if Autopilot or Anthos Service Mesh is in use).

Key tasks unique to GKE:

- **Workload Identity Federation removal**: remove `iam.gke.io/gcp-service-account` annotations; replace with service account JSON keys (hybrid period) or OpenStack application credentials.
- **Config Connector CRDs**: remove all `cnrm.cloud.google.com` CRDs; provision equivalent OpenStack resources via OpenTofu.
- **GKE Ingress to nginx-ingress**: replace `ingressClassName: gce` with `ingressClassName: nginx`; remove `ManagedCertificate`, `BackendConfig`, and `FrontendConfig` CRDs; replace with cert-manager.
- **GCE PD CSI to Cinder**: change `storageClassName` from `standard`/`standard-rwo`/`premium-rwo` to `cinder-flash`.
- **Autopilot constraints**: re-enable DaemonSets (CCM, CSI, Calico all require them), allow privileged containers, remove Autopilot-specific node tolerations.
- **Anthos Service Mesh**: uninstall ASM; install upstream Istio or Linkerd; re-apply service mesh policies using upstream CRDs.

---





### Q: What is the recommended migration path from DigitalOcean DOKS? [T1]


**Complexity: Low. Estimated timeline: 1 to 2 weeks.** DOKS is the simplest migration because it has minimal vendor lock-in.

Key tasks:

- **DO CSI to Cinder**: change `storageClassName` from `do-block-storage` to `cinder-flash`.
- **DO CCM to OpenStack CCM**: remove `do-loadbalancer-*` annotations from Services; the destination cloud controller handles standard `LoadBalancer` Services.
- **Ingress**: if already using nginx-ingress (common on DOKS), **no ingress changes are required** beyond DNS updates.
- **No pod identity to migrate**: DOKS has no IAM-for-pods mechanism; remove `DIGITALOCEAN_ACCESS_TOKEN` secret references.
- **DO Container Registry images**: mirror all images to Docker Hub, Harbor, or Quay.io before cutover; DOCR tokens expire.

---





### Q: What is the recommended migration path from Hetzner Managed Kubernetes? [T1]


**Complexity: Low-Medium. Estimated timeline: 1 to 2 weeks.** Hetzner's architecture is structurally identical to the self-managed path on Quake AI, swap the cloud provider components and rewrite the OpenTofu IaC.

Key tasks:

- **hcloud CSI to Cinder**: change `storageClassName` from `hcloud-volumes` to `cinder-flash`.
- **hcloud CCM to OpenStack CCM**: replace the `hcloud` secret in `kube-system` with the OpenStack `cloud-config` secret; remove `load-balancer.hetzner.cloud/*` annotations from Services.
- **IaC rewrite**: update OpenTofu from the `hcloud` provider to the `openstack` provider (resource concepts are similar: `hcloud_server` → `openstack_compute_instance_v2`, `hcloud_network` → `openstack_networking_network_v2`, etc.).
- **Cluster Autoscaler**: swap `--cloud-provider=hcloud` for `--cloud-provider=openstack`; reconfigure node groups to use Nova server groups.
- If using nginx-ingress (common on Hetzner). It migrates with **zero ingress changes** beyond DNS updates.

---





## Pricing, quotas, availability, and support





### Q: What regions and availability zones is the Kubernetes service available in? [T2]


The Kubernetes service runs in all three Quake AI regions: `us-east-1`, `us-east-2`, and `us-west-1`. Each cluster lives in a single region; cross-region cluster federation is not provided. Quake AI uses regions, not availability zones, so master and worker placement diversity within a region is expressed through anti-affinity [server groups](/docs/compute/concepts/server-groups), not through AZ selection. See [Regions](/docs/platform#regions) for the current region table and console URLs.

---





### Q: What is the SLA for the Kubernetes service? [T2]


To raise effective control-plane availability, run masters in an anti-affinity server group across at least three Compute instances and back up etcd on a schedule. Master nodes run in your project as Compute instances, so you own control-plane uptime rather than relying on a managed control-plane SLA from EKS, GKE, or AKS. What is covered by the platform SLA is the underlying Compute, Network, and Block Storage primitives the cluster runs on; those contractual terms are published at [rumble.cloud/legal](https://rumble.cloud/legal) and explained on the [Service Level Agreement page](/docs/account/sla).

---





### Q: How do I get help if my cluster is broken? [T2]


For most problems, start with the troubleshooting runbook:

- [Kubernetes cluster troubleshooting](/docs/operate/runbooks/kubernetes-troubleshooting): covers cluster stuck in `CREATE_IN_PROGRESS` and `kubectl` connection failures, with diagnostic commands and escalation criteria.

Before opening a support ticket, gather the following evidence:

| Item | Command |
|---|---|
| Cluster ID, name, and status | `openstack coe cluster show YOUR_CLUSTER` |
| Associated Heat stack ID and status | `openstack coe cluster show YOUR_CLUSTER -f value -c stack_id`, then `openstack stack show YOUR_STACK_ID` |
| Failed stack resources | `openstack stack resource list YOUR_STACK_ID --filter status=FAILED` |
| API reachability test | `curl -kv https://API_SERVER_IP:6443/version` |
| Console log from a stuck master | `openstack console log show YOUR_INSTANCE_NAME --lines 100` |

If the runbook does not resolve the issue, [open a support ticket](https://rumble.cloud/support) with the evidence above. [Get help](/docs/get-help) is the canonical decision tree for picking the right starting point based on the symptom, and [Support ticket evidence collection](/docs/operate/troubleshooting/support-ticket-evidence) is the full evidence checklist.

---





## See also

- [Kubernetes service overview](/docs/kubernetes): use cases, guides index, and security considerations
- [Kubernetes on Quake AI](/docs/kubernetes/concepts/kubernetes): architecture, cluster templates, resource planning, and lifecycle
- [Cloud-Native Computing on Quake AI](/docs/kubernetes/concepts/cloud-native-computing): containers vs. VMs, orchestration, CI/CD patterns
- [How to create a cluster template](/docs/kubernetes/how-to/create-cluster-template)
- [How to create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster)
- [How to manage a Kubernetes cluster](/docs/kubernetes/how-to/manage-cluster)
- [Kubernetes cluster troubleshooting](/docs/operate/runbooks/kubernetes-troubleshooting): symptom → cause → fix for cluster creation and `kubectl` failures
- [Kubernetes migration overview](/docs/kubernetes/migration): portability matrix, recommended stack, and Velero data migration
- [Migrate from AWS EKS](/docs/kubernetes/migration/migrate-from-eks)
- [Migrate from Azure AKS](/docs/kubernetes/migration/migrate-from-aks)
- [Migrate from GCP GKE](/docs/kubernetes/migration/migrate-from-gke)
- [Migrate from DigitalOcean DOKS](/docs/kubernetes/migration/migrate-from-doks)
- [Migrate from Hetzner Managed Kubernetes](/docs/kubernetes/migration/migrate-from-hetzner-k8s)
- [Coming from AWS: Kubernetes: EKS → Magnum](/resources/migration/coming-from-aws)
- [Server groups](/docs/compute/concepts/server-groups): how anti-affinity placement works at the Compute level
- [Kubernetes CLI reference](/reference/kubernetes/cli)
- [Kubernetes console reference](/reference/kubernetes/console)
