# Migrate from Kubernetes on Hetzner Cloud to Kubernetes on Quake AI

Source: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-hetzner-k8s
Markdown: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-hetzner-k8s.md

---

# Migrate from Kubernetes on Hetzner Cloud to Kubernetes on Quake AI

## Service mapping

<MigrationTable provider="hetzner" service="kubernetes" />


## 1. Overview

Kubernetes on Hetzner Cloud is self-managed: you run the control plane and worker nodes on Hetzner Cloud Servers using the hcloud cloud controller manager (`hcloud-cloud-controller-manager`) and hcloud CSI driver (`csi.hetzner.cloud`). Hetzner does not offer a native managed Kubernetes product; the typical setup is a k3s or RKE2 cluster provisioned via OpenTofu with the hcloud provider, or via community tools such as kube-hetzner or hetzner-k3s. This architecture maps closely onto the self-managed Kubernetes path on Quake AI: swap the hcloud CCM for the OpenStack CCM, swap hcloud CSI for Cinder CSI, and the rest carries over.

On Quake AI, you provision Kubernetes through [Magnum](/docs/kubernetes), the Container Infrastructure Management service. Magnum gives you a control plane and worker nodes from a curated cluster template (`openstack coe cluster create`). The us-east-1 region ships platform templates pre-wired with Calico CNI, the OpenStack cloud-provider stack (CCM + Cinder CSI), and a load-balanced API endpoint. Run `openstack coe cluster template list` for the current template catalog and Kubernetes versions, or see the [cluster templates reference](/reference/kubernetes/console/cluster-templates).

For teams that need a non-Magnum stack (custom CNI, custom kubelet flags, custom CRI, or a Kubernetes version outside the template catalog), **self-managed RKE2 or k3s on Nova instances** is documented as an alternative later in this tutorial. Hetzner users accustomed to infrastructure-as-code workflows (OpenTofu with the `hcloud` provider) can transition to the `openstack` provider with similar resource concepts in either path.

**Migration complexity: Low-Medium.** No IAM-for-pods, no proprietary CNI, and the cloud-provider swap is clean. The main effort is the OpenTofu IaC rewrite from `hcloud` to `openstack`. Expect **1-2 weeks** for a typical production migration.

## 2. Workload portability matrix

| Resource type | Portability | Notes |
|---|---|---|
| Deployment | Portable | No changes needed |
| StatefulSet (manifest) | Portable | Data migrated separately via Velero |
| DaemonSet | Portable | No changes needed |
| ConfigMap | Portable | Update any Hetzner API endpoint values |
| Secret (values) | Portable | Remove `HCLOUD_TOKEN` references |
| Service (ClusterIP, NodePort) | Portable | No changes |
| Service (LoadBalancer) | Needs adaptation | Remove `load-balancer.hetzner.cloud/*` annotations |
| Ingress | Portable | nginx-ingress is commonly used on Hetzner already |
| PersistentVolumeClaim | Needs adaptation | Change `storageClassName` from `hcloud-volumes` to `cinder-flash` |
| PersistentVolume (data) | Not portable | Migrate data via Velero filesystem backup |
| NetworkPolicy | Portable | Standard K8s NetworkPolicy spec |
| HPA / VPA / PDB | Portable | No changes |
| CronJob / Job | Portable | No changes |
| RBAC resources | Portable | No changes |
| hcloud CCM | Provider-specific | Replace with `openstack-cloud-controller-manager` |
| hcloud CSI (`csi.hetzner.cloud`) | Provider-specific | Replace with Cinder CSI |
| `hcloud` secret in kube-system | Provider-specific | Replace with `cloud-config` secret for OpenStack |

## 3. Pre-migration: export and audit

```bash
kubectl get all --all-namespaces -o yaml > cluster-export.yaml

kubectl get pvc,pv --all-namespaces -o yaml > storage.yaml

kubectl get svc --all-namespaces -o yaml | grep 'load-balancer.hetzner.cloud' > hetzner-lb-services.txt

helm list --all-namespaces > helm-releases.txt
```

**Inventory checklist:**

- [ ] All PVCs using `hcloud-volumes` StorageClass
- [ ] All Services with `load-balancer.hetzner.cloud/*` annotations
- [ ] CCM configuration: note if "networks support" is enabled (Hetzner private networking for pod routing)
- [ ] Helm releases with Hetzner-specific values
- [ ] `hcloud` secret in `kube-system` namespace

## 4. Provision Kubernetes on Quake AI

### 4.1 Primary path: Magnum

Provision a cluster from a Magnum template. The platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API.



`openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Trust delegation is not available to Keystone application credentials, so the call fails before any Magnum work begins. Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` (password auth) before running the command. See the [Kubernetes FAQ](/docs/kubernetes/faq) for the full troubleshooting flow.



```bash
openstack coe cluster template list
openstack coe cluster create production-k8s \
  --cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
  --master-count 3 \
  --node-count 3 \
  --keypair MY_KEYPAIR
```

For the full walkthrough including Console and quota guidance, see [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster). To roll your own template (custom flavors, alternate Kubernetes version), see [Create a cluster template](/docs/kubernetes/how-to/create-cluster-template).

Once the cluster reports `CREATE_COMPLETE`, fetch the kubeconfig:

```bash
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclass
```

The template pre-installs the Cinder CSI driver and the OpenStack cloud-provider stack; the Magnum path needs no manual CCM, CSI, or CNI install.

### 4.2 Alternative: self-managed RKE2 or k3s

If the Magnum template catalog does not match your version, CNI, or CRI constraints, the provisioning workflow mirrors Hetzner's self-managed approach: change the OpenTofu provider from `hcloud` to `openstack` and adjust resource types as follows:

| Hetzner resource | OpenStack equivalent |
|---|---|
| `hcloud_server` | `openstack_compute_instance_v2` |
| `hcloud_network` | `openstack_networking_network_v2` |
| `hcloud_network_subnet` | `openstack_networking_subnet_v2` |
| `hcloud_load_balancer` | Kubernetes `LoadBalancer` Service for application traffic |
| `hcloud_firewall` | `openstack_networking_secgroup_v2` |

Install RKE2 or k3s (both are popular choices on Hetzner), then deploy on the self-managed cluster:
1. `openstack-cloud-controller-manager` (replaces `hcloud-cloud-controller-manager`)
2. Cinder CSI driver (replaces `csi.hetzner.cloud`)
3. Calico (VXLAN) or Cilium as the CNI

## 5. Adapt provider-specific resources

### CSI driver: hcloud volumes to Cinder

Replace the StorageClass reference in all PVCs:

```yaml
# Before (Hetzner)
storageClassName: hcloud-volumes

# After (Quake AI)
storageClassName: cinder-flash
```

Hetzner volumes are straightforward RWO block storage with no topology constraints, a direct 1:1 replacement with Cinder.

### Ingress controller

If using nginx-ingress (common on Hetzner), it migrates with **zero changes** beyond DNS updates.

Remove Hetzner-specific load balancer annotations from Services:

```yaml
# Remove these Hetzner-specific annotations
load-balancer.hetzner.cloud/name
load-balancer.hetzner.cloud/type
load-balancer.hetzner.cloud/location
```

The cloud controller assigns a public endpoint to a Kubernetes Service with `type: LoadBalancer` without requiring provider-specific annotations.

### Cloud controller secret

Replace the `hcloud` secret in `kube-system` with the OpenStack `cloud-config` secret:

```bash
kubectl delete secret hcloud -n kube-system

kubectl create secret generic cloud-config -n kube-system \
  --from-file=cloud.conf=/path/to/cloud.conf
```

### Pod identity

Hetzner has no pod-level identity mechanism: applications authenticate to the Hetzner API using long-lived tokens in K8s Secrets.

- Remove any `HCLOUD_TOKEN` secret references
- For applications calling Quake AI/OpenStack APIs, create [application credentials](/docs/identity/how-to/create-application-credential) and store them in K8s Secrets
- All RBAC resources migrate without changes

### Helm chart values

Update Helm values files that contain Hetzner-specific cloud provider configuration. The structure is similar; cloud provider config section and StorageClass references are the primary changes.

## 6. Apply and validate

**Migrate data with Velero:**

Hetzner Object Storage is S3-compatible, so the Velero AWS plugin works against it once you set `s3ForcePathStyle=true` and point `s3Url` at your bucket's endpoint. Find the endpoint URL in the Hetzner Console, replace `REGION` with your bucket's region, and set `VELERO_AWS_PLUGIN_VERSION` to the [velero-plugin-for-aws release that matches your Velero version](https://github.com/vmware-tanzu/velero-plugin-for-aws#compatibility).

```bash
# On Hetzner: install Velero (Hetzner Object Storage is S3-compatible)
velero install \
  --provider aws \
  --plugins velero/velero-plugin-for-aws:VELERO_AWS_PLUGIN_VERSION \
  --bucket hetzner-migration \
  --use-node-agent \
  --default-volumes-to-fs-backup \
  --backup-location-config \
    region=REGION,s3ForcePathStyle=true,s3Url=https://REGION.your-objectstorage.com

velero backup create hetzner-full --include-namespaces production

# On Quake AI: restore with StorageClass remapping
kubectl apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/plugin-config: ""
    velero.io/change-storage-class: RestoreItemAction
data:
  hcloud-volumes: cinder-flash
EOF

velero restore create --from-backup hetzner-full --restore-volumes=true
```

**Deploy and verify:**

```bash
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces
```

## 7. Observability setup

Hetzner provides no managed monitoring. Most Hetzner users already run self-managed Prometheus + Grafana; these are **fully portable**. Only PVC StorageClass references need updating to `cinder-flash`.

If starting fresh:

```bash
helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set promtail.enabled=true \
  --set loki.persistence.storageClassName=cinder-flash
```

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the `kube-prometheus-stack` chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin `--version 72.9.1`, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the `--version` flag to install the current chart.

## 8. Auto-scaling alternatives

| Hetzner feature | Quake AI equivalent |
|---|---|
| Cluster Autoscaler (hcloud provider) | Cluster Autoscaler with OpenStack provider |
| HPA | Fully portable |
| VPA | Fully portable |

Swap the Cluster Autoscaler `--cloud-provider=hcloud` flag for `--cloud-provider=openstack` and configure node groups using the OpenStack Cluster Autoscaler's MachineSet or Cluster API Provider OpenStack (CAPO) MachineDeployment model. Nova server groups are placement primitives (for affinity and anti-affinity), not Cluster Autoscaler node group definitions. HPA and VPA are portable without changes.

## 9. Validation checklist

- [ ] All pods in Running or Completed state
- [ ] PVCs bound to Cinder volumes with correct data
- [ ] Ingress routes working through the ingress controller's public `LoadBalancer` Service
- [ ] DNS resolves to new floating IPs
- [ ] Prometheus scraping all targets
- [ ] No `load-balancer.hetzner.cloud/*` annotations remaining
- [ ] No `HCLOUD_TOKEN` secrets remaining
- [ ] `cloud-config` secret correctly configured in `kube-system`
- [ ] StatefulSet data integrity verified
- [ ] Application health checks passing

## 10. Provider-specific gotchas

| Gotcha | Impact | Mitigation |
|---|---|---|
| hcloud CCM "networks support" uses Hetzner private networking for pod routing | Pod IPs are Hetzner-network-based; won't work on Quake AI | Switch to VXLAN overlay CNI; pod CIDRs change |
| hcloud CSI requires `hcloud` secret in kube-system | Missing secret breaks volume provisioning after migration | Replace with `cloud-config` secret for Cinder CSI |
| Hetzner volume provisioning can take up to 10 minutes | Velero restore may time out waiting for PVCs | Increase Velero timeout: `--item-operation-timeout=30m` |
| OpenTofu `hcloud` provider to `openstack` provider | Different resource types and arguments | Rewrite IaC; resource concepts are similar but not identical |
| Hetzner robot / dedicated server clusters | Uses different CCM/CSI for bare metal | Only applies to `--cloud-provider=hcloud-robot` clusters; not typical |

## See also

- [Kubernetes migration overview](/docs/kubernetes/migration): all provider guides and portability matrix
- [Migrate from Hetzner](/resources/migration/from-hetzner): cross-service Hetzner migration hub
- [Network migration from Hetzner Networks](/docs/network/migration/migrate-from-hetzner-networks): networking-specific migration
