# Migrate from GCP GKE to Kubernetes on Quake AI

Source: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-gke
Markdown: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-gke.md

---

# Migrate from GCP GKE to Kubernetes on Quake AI

## Service mapping

<MigrationTable provider="gcp" service="kubernetes" />


## 1. Overview

Google Kubernetes Engine (GKE) provides a managed K8s control plane with deep GCP integration: Workload Identity Federation for pod-level IAM, Config Connector for provisioning GCP resources via K8s manifests, GCE Persistent Disk CSI, GCLB-based ingress (GKE Ingress), Autopilot mode with fully managed node provisioning, and GKE Dataplane V2 (Cilium-based networking).

On Quake AI, you run and maintain Kubernetes yourself on two supported paths. You can provision a cluster through [Magnum](/docs/kubernetes), the Container Infrastructure Management service, which builds a control plane and worker nodes from a curated cluster template (`openstack coe cluster create`), or you can run self-managed Kubernetes on Nova instances (kubeadm, k3s, or RKE2) with the OpenStack cloud-provider stack (CCM and Cinder CSI). Self-managed gives you control over the exact Kubernetes version, CNI, and CRI, which helps when the template catalog does not cover the version your source GKE cluster runs. Run `openstack coe cluster template list` to see the versions the catalog currently offers.

**Self-managed RKE2 or k3s on Nova instances** gives you full control over the CNI, kubelet flags, CRI, and Kubernetes version. It is the recommended path for migrated workloads and is documented later in this tutorial.

**Migration complexity: Medium-High.** Workload Identity Federation, Config Connector CRDs, GKE-specific Ingress extensions (BackendConfig, FrontendConfig, ManagedCertificate), and Autopilot constraints are the primary drivers. Expect **3-6 weeks** for a production migration, longer if Autopilot or Anthos Service Mesh is in use.

## 2. Workload portability matrix

| Resource type | Portability | Notes |
|---|---|---|
| Deployment | Portable | Remove Autopilot-specific constraints if applicable |
| StatefulSet (manifest) | Portable | Data migrated separately via Velero |
| DaemonSet | Portable | Autopilot restricts DaemonSets only in GKE-managed namespaces (`kube-system`); user DaemonSets are portable |
| ConfigMap | Portable | Update any GCP endpoint values |
| Secret (values) | Portable | Re-encode if sourced from Secret Manager |
| Service (ClusterIP, NodePort) | Portable | No changes |
| Service (LoadBalancer) | Needs adaptation | Remove GKE-specific annotations; the cloud controller assigns the public endpoint |
| Ingress | Needs adaptation | Replace GKE Ingress (`gce` class) with nginx-ingress |
| PersistentVolumeClaim | Needs adaptation | Change `storageClassName` from `standard`/`premium-rwo` to `cinder-flash` |
| PersistentVolume (data) | Not portable | Migrate data via Velero filesystem backup |
| NetworkPolicy | Portable | Requires Calico/Cilium on Quake AI |
| HPA / VPA / PDB | Portable | No changes |
| CronJob / Job | Portable | No changes |
| ServiceAccount | Needs adaptation | Remove `iam.gke.io/gcp-service-account` annotation |
| RBAC resources | Portable | No changes |
| GCE PD CSI (`pd.csi.storage.gke.io`) | Provider-specific | Replace with Cinder CSI |
| GKE Ingress (GCLB) | Provider-specific | Replace with ingress-nginx or Traefik and a `LoadBalancer` Service |
| Config Connector CRDs | Provider-specific | Remove; provision OpenStack resources via OpenTofu |
| ManagedCertificate CRD | Provider-specific | Replace with cert-manager + Let's Encrypt |
| BackendConfig / FrontendConfig CRDs | Provider-specific | Remove; use nginx-ingress annotations |
| Anthos Service Mesh | Provider-specific | Replace with self-managed Istio or Linkerd |
| GKE Dataplane V2 (Cilium via eBPF) | Provider-specific | Install standard open-source Cilium. GKE Dataplane V2 is a modified Cilium build and does not support upstream CiliumNetworkPolicy CRDs on older versions |

## 3. Pre-migration: export and audit

```bash
kubectl get all --all-namespaces -o yaml > cluster-export.yaml

kubectl get pvc --all-namespaces -o yaml > pvcs.yaml

kubectl get sa --all-namespaces -o yaml | grep 'iam.gke.io' > workload-id-accounts.txt

# Workload Identity Federation for GKE (renamed 2024) also supports direct IAM
# principal bindings that carry no ServiceAccount annotation; list them from GCP IAM:
gcloud projects get-iam-policy PROJECT_ID --format=json > gcp-iam-policy.json

kubectl get managedcertificates --all-namespaces -o yaml > managed-certs.yaml

kubectl api-resources | grep cnrm.cloud.google.com > config-connector-crds.txt

helm list --all-namespaces > helm-releases.txt
```

**Inventory checklist:**

- [ ] All PVCs using `standard`, `standard-rwo`, or `premium-rwo` StorageClasses
- [ ] All ServiceAccounts with the `iam.gke.io/gcp-service-account` annotation, plus direct IAM principal bindings (Workload Identity Federation for GKE) that carry no annotation
- [ ] All Config Connector CRDs (`cnrm.cloud.google.com` resources) and which GCP resources they manage
- [ ] All GKE ManagedCertificate resources
- [ ] All BackendConfig / FrontendConfig resources
- [ ] Anthos Service Mesh configuration (VirtualService, DestinationRule)
- [ ] Autopilot constraints: DaemonSets, privileged containers, host network
- [ ] Artifact Registry (`REGION-docker.pkg.dev`) image references
- [ ] Regional PD (multi-zone) PVCs
- [ ] Filestore (RWX) PVCs

## 4. Provision Kubernetes on Quake AI

This guide documents both provisioning paths: Magnum in section 4.1 and self-managed Kubernetes in section 4.2. Both are supported; choose based on how closely you need to match your source cluster's Kubernetes version and CNI.

### 4.1 Magnum

To provision a cluster from a Magnum template, the platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API. The Kubernetes version comes from the template catalog; if you need a version or CNI it does not cover, use the self-managed path in section 4.2.


**Magnum cluster creation requires a password-scoped session.** `openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Trust delegation is not available to Keystone application credentials, so the call fails before any Magnum work begins. Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` (password auth) before running the command. The Cloud Console wizard works from any logged-in user session because it uses the user's password-scoped token. See [the Kubernetes FAQ](/docs/kubernetes/faq) for the full troubleshooting flow.


```bash
openstack coe cluster template list
openstack coe cluster create production-k8s \
  --cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
  --master-count 3 \
  --node-count 3 \
  --keypair MY_KEYPAIR
```

For the full walkthrough including Console and quota guidance, see [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster). To roll your own template (custom flavors, alternate Kubernetes version), see [Create a cluster template](/docs/kubernetes/how-to/create-cluster-template).

Once the cluster reports `CREATE_COMPLETE`, fetch the kubeconfig:

```bash
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclass
```

The Cinder CSI driver and the OpenStack cloud-provider stack are pre-installed by the template; no manual CCM, CSI, or CNI install is required for the Magnum path.

### 4.2 Recommended path: self-managed RKE2 or k3s

Build the cluster yourself on Nova instances for full control over the Kubernetes version, CNI, and CRI. This is the recommended path for migrated workloads. Use OpenTofu with the `openstack` provider to create the infrastructure, then install Kubernetes.

**Infrastructure (OpenTofu):**
- Neutron network + subnet
- Security groups (K8s API 6443, NodePort 30000-32767, inter-node)
- Nova instances: 3 control plane + N workers
- Stable endpoint for the Kubernetes API

**Install RKE2 (one supported self-managed distribution):**

```bash
curl -sfL https://get.rke2.io | INSTALL_RKE2_TYPE=server sh -
systemctl enable --now rke2-server
```

Deploy the OpenStack cloud-provider stack on the self-managed cluster:
1. `openstack-cloud-controller-manager` with [application credentials](/docs/identity/how-to/create-application-credential)
2. Cinder CSI driver with `cinder-flash` StorageClass
3. Calico (VXLAN) or Cilium as the CNI

## 5. Adapt provider-specific resources

### CSI driver: GCE persistent disk to Cinder

Replace all StorageClass references:

```yaml
# Before (GKE): standard, standard-rwo, or premium-rwo
storageClassName: standard

# After (Quake AI)
storageClassName: cinder-flash
```

**Regional PDs:** GKE supports multi-zone persistent disks for HA StatefulSets. Cinder volumes on Quake AI are single-AZ. For HA stateful workloads, implement replication at the application layer (database streaming replication, Redis Sentinel) rather than relying on multi-AZ disk replication.

**Filestore (RWX):** GKE Filestore provides ReadWriteMany access. Options on Quake AI:
- NFS server provisioner (dedicated Nova VM)
- Rook-Ceph with CephFS (production-grade RWX)
- Application restructuring to use RWO volumes

### Ingress controller: GKE ingress to nginx-ingress

GKE Ingress uses Google's HTTP Load Balancer (GCLB) with GKE-specific CRDs. Replace with nginx-ingress:

```yaml
# Before (GKE Ingress)
annotations:
  kubernetes.io/ingress.class: "gce"
  kubernetes.io/ingress.global-static-ip-name: "my-static-ip"

# After (Quake AI nginx-ingress)
annotations:
  kubernetes.io/ingress.class: nginx
```

**Remove GKE-specific CRDs:**
- `ManagedCertificate`: replace with cert-manager `Certificate` resources using Let's Encrypt
- `BackendConfig`: remove; configure health checks and timeouts via nginx-ingress annotations
- `FrontendConfig`: remove; configure redirects via nginx-ingress annotations

### Pod identity: workload identity to application credentials

GKE renamed this feature to Workload Identity Federation for GKE in 2024. It has two approaches: ServiceAccount impersonation, which annotates a ServiceAccount with `iam.gke.io/gcp-service-account` (the older approach), and direct IAM principal binding, which grants permissions to Kubernetes resources through a principal identifier and carries no ServiceAccount annotation. Both exchange Kubernetes tokens for GCP access through Google's security token service, and both are GCP-specific and cannot work outside GKE. Audit both paths: check ServiceAccount annotations and the project IAM policy (`gcloud projects get-iam-policy PROJECT_ID`) for GKE principal bindings.

**For applications migrating to OpenStack services (GCS to Swift, Pub/Sub to RabbitMQ):** Use [application credentials](/docs/identity/how-to/create-application-credential) in K8s Secrets.

**For applications that still need GCP access (hybrid period):** Use service account JSON keys:

```bash
kubectl create secret generic gcp-sa-key \
  --from-file=key.json=/path/to/service-account.json
```

Mount as a volume and set `GOOGLE_APPLICATION_CREDENTIALS` to the mount path.

Remove the Workload Identity annotation from every affected ServiceAccount:

```bash
kubectl annotate sa SA_NAME iam.gke.io/gcp-service-account- -n NAMESPACE
```

### Config connector cRDs

Config Connector creates GCP resources (BigQuery datasets, Cloud SQL instances, Pub/Sub topics) from K8s manifests. These CRDs only function in GKE clusters with GCP access.

Remove all Config Connector resources before restoring on Quake AI:

```bash
kubectl get crds | grep cnrm.cloud.google.com | awk '{print $1}' | xargs kubectl delete crd
```

Provision equivalent OpenStack resources via OpenTofu instead.

### Service type LoadBalancer

Remove GKE-specific annotations. The cloud controller assigns a public endpoint to a Kubernetes Service with `type: LoadBalancer`.

### Autopilot constraints

If migrating from GKE Autopilot:
- Deploy the CCM, CSI, and CNI DaemonSets that Autopilot blocked from `kube-system`. Autopilot restricts DaemonSets only in GKE-managed namespaces, not user-deployed DaemonSets, so these run normally on Quake AI.
- Allow privileged containers where needed
- Remove `node.kubernetes.io/workload` tolerations that were Autopilot-specific
- Remove minimum resource request constraints imposed by Autopilot

### Container images: Artifact Registry

If using Artifact Registry (`REGION-docker.pkg.dev`) with GKE-integrated authentication, mirror images to a permanent registry (Docker Hub, Harbor). Artifact Registry auth tokens tied to GKE metadata do not work from Quake AI. Google Container Registry (`gcr.io`) was shut down in March 2025; any remaining images should already have moved to Artifact Registry.

### Anthos Service mesh

If using Anthos Service Mesh (ASM), uninstall it and install upstream Istio or Linkerd. ASM CRDs differ from upstream Istio. Re-apply service mesh policies (VirtualService, DestinationRule) using upstream Istio CRDs.

## 6. Apply and validate

**Migrate data with Velero:**

```bash
# On GKE: install Velero with GCP backend
velero install \
  --provider gcp \
  --plugins velero/velero-plugin-for-gcp:v1.9.0 \
  --bucket gke-migration-backup \
  --use-node-agent \
  --default-volumes-to-fs-backup

velero backup create gke-full --include-namespaces production

# On Quake AI: restore with StorageClass remapping
kubectl apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/plugin-config: ""
    velero.io/change-storage-class: RestoreItemAction
data:
  standard: cinder-flash
  standard-rwo: cinder-flash
  premium-rwo: cinder-flash
EOF

velero restore create --from-backup gke-full --restore-volumes=true
```

**Deploy and verify:**

```bash
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces
kubectl get ingress --all-namespaces
```

For Cloud SQL databases, use managed export/import to self-hosted PostgreSQL or MySQL on Quake AI. For GCS data, use `rclone copy gs://bucket s3://swift-bucket`.

## 7. Observability setup

Replace Google Cloud Operations (Kubernetes Engine Monitoring, Cloud Logging, Cloud Trace) with a self-managed stack:

```bash
helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set promtail.enabled=true \
  --set loki.persistence.storageClassName=cinder-flash
```

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the `kube-prometheus-stack` chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin `--version 72.9.1`, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the `--version` flag to install the current chart.

For distributed tracing (replacing Cloud Trace), add Grafana Tempo.

## 8. Auto-scaling alternatives

| GKE feature | Quake AI equivalent |
|---|---|
| Autopilot (fully managed nodes, no node pools) | No equivalent. Pre-provision node pools in OpenTofu and configure the Cluster Autoscaler. Size nodes from your workload resource requests, which Autopilot sized automatically |
| Cluster Autoscaler (Standard mode) | Cluster Autoscaler with OpenStack provider |
| Node Auto-Provisioning (NAP) | Not available; pre-provision node pools of different flavors |
| VPA | Fully portable |
| HPA | Fully portable |

If migrating from Autopilot, you must explicitly size and manage worker nodes. Pre-provision node pools with OpenTofu and configure the Cluster Autoscaler with `--cloud-provider=openstack`.

## 9. Validation checklist

- [ ] All pods in Running or Completed state
- [ ] PVCs bound to Cinder volumes with correct data
- [ ] Ingress routes working through the ingress controller's public `LoadBalancer` Service
- [ ] DNS resolves to the new Service address
- [ ] cert-manager certificates issued (replacing ManagedCertificate)
- [ ] No `iam.gke.io` annotations remaining on ServiceAccounts
- [ ] No Config Connector CRDs (`cnrm.cloud.google.com`) remaining
- [ ] No BackendConfig/FrontendConfig/ManagedCertificate CRDs remaining
- [ ] Prometheus scraping all targets
- [ ] Loki ingesting logs
- [ ] StatefulSet data integrity verified
- [ ] Application health checks passing
- [ ] DaemonSets running (if migrated from Autopilot)

## 10. Provider-specific gotchas

| Gotcha | Impact | Mitigation |
|---|---|---|
| Autopilot restricts DaemonSets in `kube-system` | CCM, CSI, and CNI must run as DaemonSets | Deploy them yourself on Quake AI; user DaemonSets run normally |
| Workload Identity Federation is GCP-only | OIDC token exchange doesn't work outside GCP | Replace with service account JSON keys or migrate to OpenStack services |
| Config Connector CRDs only work in GKE | CRDs become invalid on non-GKE clusters | Remove before restore; provision equivalents via OpenTofu |
| GKE Backup (Backup for GKE) format is GCP-native | Cannot restore to Quake AI directly | Use Velero + filesystem backup instead |
| Regional PDs (multi-zone HA storage) | Cinder is single-AZ per volume | Implement application-level HA (database replication) |
| GKE Ingress creates Cloud Armor WAF rules | The destination cluster needs an application-layer WAF | Deploy ModSecurity with ingress-nginx, or place an external CDN and WAF in front of the Service address |
| `cloud.google.com/` node labels | May break nodeAffinity scheduling rules | Update nodeAffinity to use standard K8s labels |
| Anthos Service Mesh CRDs differ from upstream Istio | ASM VirtualService/DestinationRule may need adjustment | Uninstall ASM; install upstream Istio; re-apply policies |
| GKE Enterprise Config Management (GitOps) has no Quake AI equivalent | Policy-as-code and config sync stop working off GKE | Replace with Argo CD or Flux |
| GKE Enterprise Policy Controller (OPA-based) has no Quake AI equivalent | Admission policies stop enforcing off GKE | Replace with OPA Gatekeeper or Kyverno (both install via Helm) |
| Filestore (RWX) has no direct Cinder equivalent | Applications requiring shared storage need restructuring | NFS server provisioner, Rook-Ceph, or application-level coordination |

## See also

- [Kubernetes migration overview](/docs/kubernetes/migration): all provider guides and portability matrix
- [Migrate from GCP](/resources/migration/from-gcp): cross-service GCP migration hub
- [Network migration from GCP VPC](/docs/network/migration/migrate-from-gcp-vpc): networking-specific migration
- [Compute migration from GCE](/docs/compute/migration/migrate-from-gce): VM-level migration
