Migrate from GCP GKE to Kubernetes on Quake AI
Coming from another cloud?
▸Google Cloud·GKE Cluster
GKE Cluster
- GKE provides Autopilot mode with fully managed node provisioning and scaling by Google; on Quake AI you provision clusters through Magnum or self-managed Kubernetes on Nova instances and manage node scaling yourself.
- Control plane is fully managed with automatic upgrades through release channels; Quake AI requires user-provisioned and user-managed control plane nodes.
- Cluster creation uses gcloud CLI vs OpenStack CLI (openstack coe cluster create).
- Custom machine types, spot VMs, accelerators in node pools; Quake AI K8s nodes use standard Nova flavors.
Migrate from GCP GKE to Kubernetes on Quake AI
Service mapping#
Google Cloud to Quake AI
| Google Cloud service | Quake AI equivalent | Key difference |
|---|---|---|
| GKE Cluster | Kubernetes | GKE provides Autopilot mode with fully managed node provisioning and scaling by Google; on Quake AI you provision clusters through Magnum... see details |
1. Overview#
Google Kubernetes Engine (GKE) provides a managed K8s control plane with deep GCP integration: Workload Identity Federation for pod-level IAM, Config Connector for provisioning GCP resources via K8s manifests, GCE Persistent Disk CSI, GCLB-based ingress (GKE Ingress), Autopilot mode with fully managed node provisioning, and GKE Dataplane V2 (Cilium-based networking).
On Quake AI, you run and maintain Kubernetes yourself on two supported paths. You can provision a cluster through Magnum, the Container Infrastructure Management service, which builds a control plane and worker nodes from a curated cluster template (openstack coe cluster create), or you can run self-managed Kubernetes on Nova instances (kubeadm, k3s, or RKE2) with the OpenStack cloud-provider stack (CCM and Cinder CSI). Self-managed gives you control over the exact Kubernetes version, CNI, and CRI, which helps when the template catalog does not cover the version your source GKE cluster runs. Run openstack coe cluster template list to see the versions the catalog currently offers.
Self-managed RKE2 or k3s on Nova instances gives you full control over the CNI, kubelet flags, CRI, and Kubernetes version. It is the recommended path for migrated workloads and is documented later in this tutorial.
Migration complexity: Medium-High. Workload Identity Federation, Config Connector CRDs, GKE-specific Ingress extensions (BackendConfig, FrontendConfig, ManagedCertificate), and Autopilot constraints are the primary drivers. Expect 3-6 weeks for a production migration, longer if Autopilot or Anthos Service Mesh is in use.
2. Workload portability matrix#
| Resource type | Portability | Notes |
|---|---|---|
| Deployment | Portable | Remove Autopilot-specific constraints if applicable |
| StatefulSet (manifest) | Portable | Data migrated separately via Velero |
| DaemonSet | Portable | Autopilot restricts DaemonSets only in GKE-managed namespaces (kube-system); user DaemonSets are portable |
| ConfigMap | Portable | Update any GCP endpoint values |
| Secret (values) | Portable | Re-encode if sourced from Secret Manager |
| Service (ClusterIP, NodePort) | Portable | No changes |
| Service (LoadBalancer) | Needs adaptation | Remove GKE-specific annotations; the cloud controller assigns the public endpoint |
| Ingress | Needs adaptation | Replace GKE Ingress (gce class) with nginx-ingress |
| PersistentVolumeClaim | Needs adaptation | Change storageClassName from standard/premium-rwo to cinder-flash |
| PersistentVolume (data) | Not portable | Migrate data via Velero filesystem backup |
| NetworkPolicy | Portable | Requires Calico/Cilium on Quake AI |
| HPA / VPA / PDB | Portable | No changes |
| CronJob / Job | Portable | No changes |
| ServiceAccount | Needs adaptation | Remove iam.gke.io/gcp-service-account annotation |
| RBAC resources | Portable | No changes |
GCE PD CSI (pd.csi.storage.gke.io) | Provider-specific | Replace with Cinder CSI |
| GKE Ingress (GCLB) | Provider-specific | Replace with ingress-nginx or Traefik and a LoadBalancer Service |
| Config Connector CRDs | Provider-specific | Remove; provision OpenStack resources via OpenTofu |
| ManagedCertificate CRD | Provider-specific | Replace with cert-manager + Let's Encrypt |
| BackendConfig / FrontendConfig CRDs | Provider-specific | Remove; use nginx-ingress annotations |
| Anthos Service Mesh | Provider-specific | Replace with self-managed Istio or Linkerd |
| GKE Dataplane V2 (Cilium via eBPF) | Provider-specific | Install standard open-source Cilium. GKE Dataplane V2 is a modified Cilium build and does not support upstream CiliumNetworkPolicy CRDs on older versions |
3. Pre-migration: export and audit#
kubectl get all --all-namespaces -o yaml > cluster-export.yaml
kubectl get pvc --all-namespaces -o yaml > pvcs.yaml
kubectl get sa --all-namespaces -o yaml | grep 'iam.gke.io' > workload-id-accounts.txt
# Workload Identity Federation for GKE (renamed 2024) also supports direct IAM
# principal bindings that carry no ServiceAccount annotation; list them from GCP IAM:
gcloud projects get-iam-policy PROJECT_ID --format=json > gcp-iam-policy.json
kubectl get managedcertificates --all-namespaces -o yaml > managed-certs.yaml
kubectl api-resources | grep cnrm.cloud.google.com > config-connector-crds.txt
helm list --all-namespaces > helm-releases.txtInventory checklist:
- All PVCs using
standard,standard-rwo, orpremium-rwoStorageClasses - All ServiceAccounts with the
iam.gke.io/gcp-service-accountannotation, plus direct IAM principal bindings (Workload Identity Federation for GKE) that carry no annotation - All Config Connector CRDs (
cnrm.cloud.google.comresources) and which GCP resources they manage - All GKE ManagedCertificate resources
- All BackendConfig / FrontendConfig resources
- Anthos Service Mesh configuration (VirtualService, DestinationRule)
- Autopilot constraints: DaemonSets, privileged containers, host network
- Artifact Registry (
REGION-docker.pkg.dev) image references - Regional PD (multi-zone) PVCs
- Filestore (RWX) PVCs
4. Provision Kubernetes on Quake AI#
This guide documents both provisioning paths: Magnum in section 4.1 and self-managed Kubernetes in section 4.2. Both are supported; choose based on how closely you need to match your source cluster's Kubernetes version and CNI.
4.1 Magnum#
To provision a cluster from a Magnum template, the platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API. The Kubernetes version comes from the template catalog; if you need a version or CNI it does not cover, use the self-managed path in section 4.2.
openstack coe cluster template list
openstack coe cluster create production-k8s \
--cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
--master-count 3 \
--node-count 3 \
--keypair MY_KEYPAIRFor the full walkthrough including Console and quota guidance, see Create a Kubernetes cluster. To roll your own template (custom flavors, alternate Kubernetes version), see Create a cluster template.
Once the cluster reports CREATE_COMPLETE, fetch the kubeconfig:
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclassThe Cinder CSI driver and the OpenStack cloud-provider stack are pre-installed by the template; no manual CCM, CSI, or CNI install is required for the Magnum path.
4.2 Recommended path: self-managed RKE2 or k3s#
Build the cluster yourself on Nova instances for full control over the Kubernetes version, CNI, and CRI. This is the recommended path for migrated workloads. Use OpenTofu with the openstack provider to create the infrastructure, then install Kubernetes.
Infrastructure (OpenTofu):
- Neutron network + subnet
- Security groups (K8s API 6443, NodePort 30000-32767, inter-node)
- Nova instances: 3 control plane + N workers
- Stable endpoint for the Kubernetes API
Install RKE2 (one supported self-managed distribution):
curl -sfL https://get.rke2.io | INSTALL_RKE2_TYPE=server sh -
systemctl enable --now rke2-serverDeploy the OpenStack cloud-provider stack on the self-managed cluster:
openstack-cloud-controller-managerwith application credentials- Cinder CSI driver with
cinder-flashStorageClass - Calico (VXLAN) or Cilium as the CNI
5. Adapt provider-specific resources#
CSI driver: GCE persistent disk to Cinder#
Replace all StorageClass references:
# Before (GKE): standard, standard-rwo, or premium-rwo
storageClassName: standard
# After (Quake AI)
storageClassName: cinder-flashRegional PDs: GKE supports multi-zone persistent disks for HA StatefulSets. Cinder volumes on Quake AI are single-AZ. For HA stateful workloads, implement replication at the application layer (database streaming replication, Redis Sentinel) rather than relying on multi-AZ disk replication.
Filestore (RWX): GKE Filestore provides ReadWriteMany access. Options on Quake AI:
- NFS server provisioner (dedicated Nova VM)
- Rook-Ceph with CephFS (production-grade RWX)
- Application restructuring to use RWO volumes
Ingress controller: GKE ingress to nginx-ingress#
GKE Ingress uses Google's HTTP Load Balancer (GCLB) with GKE-specific CRDs. Replace with nginx-ingress:
# Before (GKE Ingress)
annotations:
kubernetes.io/ingress.class: "gce"
kubernetes.io/ingress.global-static-ip-name: "my-static-ip"
# After (Quake AI nginx-ingress)
annotations:
kubernetes.io/ingress.class: nginxRemove GKE-specific CRDs:
ManagedCertificate: replace with cert-managerCertificateresources using Let's EncryptBackendConfig: remove; configure health checks and timeouts via nginx-ingress annotationsFrontendConfig: remove; configure redirects via nginx-ingress annotations
Pod identity: workload identity to application credentials#
GKE renamed this feature to Workload Identity Federation for GKE in 2024. It has two approaches: ServiceAccount impersonation, which annotates a ServiceAccount with iam.gke.io/gcp-service-account (the older approach), and direct IAM principal binding, which grants permissions to Kubernetes resources through a principal identifier and carries no ServiceAccount annotation. Both exchange Kubernetes tokens for GCP access through Google's security token service, and both are GCP-specific and cannot work outside GKE. Audit both paths: check ServiceAccount annotations and the project IAM policy (gcloud projects get-iam-policy PROJECT_ID) for GKE principal bindings.
For applications migrating to OpenStack services (GCS to Swift, Pub/Sub to RabbitMQ): Use application credentials in K8s Secrets.
For applications that still need GCP access (hybrid period): Use service account JSON keys:
kubectl create secret generic gcp-sa-key \
--from-file=key.json=/path/to/service-account.jsonMount as a volume and set GOOGLE_APPLICATION_CREDENTIALS to the mount path.
Remove the Workload Identity annotation from every affected ServiceAccount:
kubectl annotate sa SA_NAME iam.gke.io/gcp-service-account- -n NAMESPACEConfig connector cRDs#
Config Connector creates GCP resources (BigQuery datasets, Cloud SQL instances, Pub/Sub topics) from K8s manifests. These CRDs only function in GKE clusters with GCP access.
Remove all Config Connector resources before restoring on Quake AI:
kubectl get crds | grep cnrm.cloud.google.com | awk '{print $1}' | xargs kubectl delete crdProvision equivalent OpenStack resources via OpenTofu instead.
Service type LoadBalancer#
Remove GKE-specific annotations. The cloud controller assigns a public endpoint to a Kubernetes Service with type: LoadBalancer.
Autopilot constraints#
If migrating from GKE Autopilot:
- Deploy the CCM, CSI, and CNI DaemonSets that Autopilot blocked from
kube-system. Autopilot restricts DaemonSets only in GKE-managed namespaces, not user-deployed DaemonSets, so these run normally on Quake AI. - Allow privileged containers where needed
- Remove
node.kubernetes.io/workloadtolerations that were Autopilot-specific - Remove minimum resource request constraints imposed by Autopilot
Container images: Artifact Registry#
If using Artifact Registry (REGION-docker.pkg.dev) with GKE-integrated authentication, mirror images to a permanent registry (Docker Hub, Harbor). Artifact Registry auth tokens tied to GKE metadata do not work from Quake AI. Google Container Registry (gcr.io) was shut down in March 2025; any remaining images should already have moved to Artifact Registry.
Anthos Service mesh#
If using Anthos Service Mesh (ASM), uninstall it and install upstream Istio or Linkerd. ASM CRDs differ from upstream Istio. Re-apply service mesh policies (VirtualService, DestinationRule) using upstream Istio CRDs.
6. Apply and validate#
Migrate data with Velero:
# On GKE: install Velero with GCP backend
velero install \
--provider gcp \
--plugins velero/velero-plugin-for-gcp:v1.9.0 \
--bucket gke-migration-backup \
--use-node-agent \
--default-volumes-to-fs-backup
velero backup create gke-full --include-namespaces production
# On Quake AI: restore with StorageClass remapping
kubectl apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
name: change-storage-class-config
namespace: velero
labels:
velero.io/plugin-config: ""
velero.io/change-storage-class: RestoreItemAction
data:
standard: cinder-flash
standard-rwo: cinder-flash
premium-rwo: cinder-flash
EOF
velero restore create --from-backup gke-full --restore-volumes=trueDeploy and verify:
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces
kubectl get ingress --all-namespacesFor Cloud SQL databases, use managed export/import to self-hosted PostgreSQL or MySQL on Quake AI. For GCS data, use rclone copy gs://bucket s3://swift-bucket.
7. Observability setup#
Replace Google Cloud Operations (Kubernetes Engine Monitoring, Cloud Logging, Cloud Trace) with a self-managed stack:
helm upgrade --install kube-prometheus-stack \
prometheus-community/kube-prometheus-stack \
--version 72.9.1 \
--namespace monitoring --create-namespace
helm upgrade --install loki grafana/loki-stack \
--namespace monitoring \
--set promtail.enabled=true \
--set loki.persistence.storageClassName=cinder-flashMatch the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the kube-prometheus-stack chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin --version 72.9.1, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the --version flag to install the current chart.
For distributed tracing (replacing Cloud Trace), add Grafana Tempo.
8. Auto-scaling alternatives#
| GKE feature | Quake AI equivalent |
|---|---|
| Autopilot (fully managed nodes, no node pools) | No equivalent. Pre-provision node pools in OpenTofu and configure the Cluster Autoscaler. Size nodes from your workload resource requests, which Autopilot sized automatically |
| Cluster Autoscaler (Standard mode) | Cluster Autoscaler with OpenStack provider |
| Node Auto-Provisioning (NAP) | Not available; pre-provision node pools of different flavors |
| VPA | Fully portable |
| HPA | Fully portable |
If migrating from Autopilot, you must explicitly size and manage worker nodes. Pre-provision node pools with OpenTofu and configure the Cluster Autoscaler with --cloud-provider=openstack.
9. Validation checklist#
- All pods in Running or Completed state
- PVCs bound to Cinder volumes with correct data
- Ingress routes working through the ingress controller's public
LoadBalancerService - DNS resolves to the new Service address
- cert-manager certificates issued (replacing ManagedCertificate)
- No
iam.gke.ioannotations remaining on ServiceAccounts - No Config Connector CRDs (
cnrm.cloud.google.com) remaining - No BackendConfig/FrontendConfig/ManagedCertificate CRDs remaining
- Prometheus scraping all targets
- Loki ingesting logs
- StatefulSet data integrity verified
- Application health checks passing
- DaemonSets running (if migrated from Autopilot)
10. Provider-specific gotchas#
| Gotcha | Impact | Mitigation |
|---|---|---|
Autopilot restricts DaemonSets in kube-system | CCM, CSI, and CNI must run as DaemonSets | Deploy them yourself on Quake AI; user DaemonSets run normally |
| Workload Identity Federation is GCP-only | OIDC token exchange doesn't work outside GCP | Replace with service account JSON keys or migrate to OpenStack services |
| Config Connector CRDs only work in GKE | CRDs become invalid on non-GKE clusters | Remove before restore; provision equivalents via OpenTofu |
| GKE Backup (Backup for GKE) format is GCP-native | Cannot restore to Quake AI directly | Use Velero + filesystem backup instead |
| Regional PDs (multi-zone HA storage) | Cinder is single-AZ per volume | Implement application-level HA (database replication) |
| GKE Ingress creates Cloud Armor WAF rules | The destination cluster needs an application-layer WAF | Deploy ModSecurity with ingress-nginx, or place an external CDN and WAF in front of the Service address |
cloud.google.com/ node labels | May break nodeAffinity scheduling rules | Update nodeAffinity to use standard K8s labels |
| Anthos Service Mesh CRDs differ from upstream Istio | ASM VirtualService/DestinationRule may need adjustment | Uninstall ASM; install upstream Istio; re-apply policies |
| GKE Enterprise Config Management (GitOps) has no Quake AI equivalent | Policy-as-code and config sync stop working off GKE | Replace with Argo CD or Flux |
| GKE Enterprise Policy Controller (OPA-based) has no Quake AI equivalent | Admission policies stop enforcing off GKE | Replace with OPA Gatekeeper or Kyverno (both install via Helm) |
| Filestore (RWX) has no direct Cinder equivalent | Applications requiring shared storage need restructuring | NFS server provisioner, Rook-Ceph, or application-level coordination |
See also#
- Kubernetes migration overview: all provider guides and portability matrix
- Migrate from GCP: cross-service GCP migration hub
- Network migration from GCP VPC: networking-specific migration
- Compute migration from GCE: VM-level migration
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
Comparisons to third-party providers in this material reflect publicly documented behavior as of the validation date below. Pricing, quotas, service limits, and feature availability change frequently on every cloud. Verify provider-specific claims against the provider's own current documentation before relying on them for a procurement, architecture, or migration decision.
For the full policy, see Usage Guidelines.
Last validated: 22.06.2026
Quick answers
See Also
Kubernetes platforms
Related
Migrate from Azure AKS to Kubernetes on Quake AI
Shares: Kubernetes, Containers
Migrate from DigitalOcean DOKS to Kubernetes on Quake AI
Shares: Kubernetes, Containers
Migrate from AWS EKS to Kubernetes on Quake AI
Shares: Kubernetes, Containers
Migrate from Kubernetes on Hetzner Cloud to Kubernetes on Quake AI
Shares: Kubernetes, Containers