Skip to content

Migrate from GCP GKE to Kubernetes on Quake AI

Migration · Updated Jun 2026

Coming from another cloud?

▸Google Cloud·GKE Cluster

GKE Clusterhigh

  • GKE provides Autopilot mode with fully managed node provisioning and scaling by Google; on Quake AI you provision clusters through Magnum or self-managed Kubernetes on Nova instances and manage node scaling yourself.
  • Control plane is fully managed with automatic upgrades through release channels; Quake AI requires user-provisioned and user-managed control plane nodes.
  • Cluster creation uses gcloud CLI vs OpenStack CLI (openstack coe cluster create).
  • Custom machine types, spot VMs, accelerators in node pools; Quake AI K8s nodes use standard Nova flavors.
Google Cloud docs ↗

Migrate from GCP GKE to Kubernetes on Quake AI

Service mapping#

Google Cloud to Quake AI

Google Cloud serviceQuake AI equivalentKey difference
GKE ClusterKubernetesGKE provides Autopilot mode with fully managed node provisioning and scaling by Google; on Quake AI you provision clusters through Magnum... see details

1. Overview#

Google Kubernetes Engine (GKE) provides a managed K8s control plane with deep GCP integration: Workload Identity Federation for pod-level IAM, Config Connector for provisioning GCP resources via K8s manifests, GCE Persistent Disk CSI, GCLB-based ingress (GKE Ingress), Autopilot mode with fully managed node provisioning, and GKE Dataplane V2 (Cilium-based networking).

On Quake AI, you run and maintain Kubernetes yourself on two supported paths. You can provision a cluster through Magnum, the Container Infrastructure Management service, which builds a control plane and worker nodes from a curated cluster template (openstack coe cluster create), or you can run self-managed Kubernetes on Nova instances (kubeadm, k3s, or RKE2) with the OpenStack cloud-provider stack (CCM and Cinder CSI). Self-managed gives you control over the exact Kubernetes version, CNI, and CRI, which helps when the template catalog does not cover the version your source GKE cluster runs. Run openstack coe cluster template list to see the versions the catalog currently offers.

Self-managed RKE2 or k3s on Nova instances gives you full control over the CNI, kubelet flags, CRI, and Kubernetes version. It is the recommended path for migrated workloads and is documented later in this tutorial.

Migration complexity: Medium-High. Workload Identity Federation, Config Connector CRDs, GKE-specific Ingress extensions (BackendConfig, FrontendConfig, ManagedCertificate), and Autopilot constraints are the primary drivers. Expect 3-6 weeks for a production migration, longer if Autopilot or Anthos Service Mesh is in use.

2. Workload portability matrix#

Resource typePortabilityNotes
DeploymentPortableRemove Autopilot-specific constraints if applicable
StatefulSet (manifest)PortableData migrated separately via Velero
DaemonSetPortableAutopilot restricts DaemonSets only in GKE-managed namespaces (kube-system); user DaemonSets are portable
ConfigMapPortableUpdate any GCP endpoint values
Secret (values)PortableRe-encode if sourced from Secret Manager
Service (ClusterIP, NodePort)PortableNo changes
Service (LoadBalancer)Needs adaptationRemove GKE-specific annotations; the cloud controller assigns the public endpoint
IngressNeeds adaptationReplace GKE Ingress (gce class) with nginx-ingress
PersistentVolumeClaimNeeds adaptationChange storageClassName from standard/premium-rwo to cinder-flash
PersistentVolume (data)Not portableMigrate data via Velero filesystem backup
NetworkPolicyPortableRequires Calico/Cilium on Quake AI
HPA / VPA / PDBPortableNo changes
CronJob / JobPortableNo changes
ServiceAccountNeeds adaptationRemove iam.gke.io/gcp-service-account annotation
RBAC resourcesPortableNo changes
GCE PD CSI (pd.csi.storage.gke.io)Provider-specificReplace with Cinder CSI
GKE Ingress (GCLB)Provider-specificReplace with ingress-nginx or Traefik and a LoadBalancer Service
Config Connector CRDsProvider-specificRemove; provision OpenStack resources via OpenTofu
ManagedCertificate CRDProvider-specificReplace with cert-manager + Let's Encrypt
BackendConfig / FrontendConfig CRDsProvider-specificRemove; use nginx-ingress annotations
Anthos Service MeshProvider-specificReplace with self-managed Istio or Linkerd
GKE Dataplane V2 (Cilium via eBPF)Provider-specificInstall standard open-source Cilium. GKE Dataplane V2 is a modified Cilium build and does not support upstream CiliumNetworkPolicy CRDs on older versions

3. Pre-migration: export and audit#

bash
kubectl get all --all-namespaces -o yaml > cluster-export.yaml

kubectl get pvc --all-namespaces -o yaml > pvcs.yaml

kubectl get sa --all-namespaces -o yaml | grep 'iam.gke.io' > workload-id-accounts.txt

# Workload Identity Federation for GKE (renamed 2024) also supports direct IAM
# principal bindings that carry no ServiceAccount annotation; list them from GCP IAM:
gcloud projects get-iam-policy PROJECT_ID --format=json > gcp-iam-policy.json

kubectl get managedcertificates --all-namespaces -o yaml > managed-certs.yaml

kubectl api-resources | grep cnrm.cloud.google.com > config-connector-crds.txt

helm list --all-namespaces > helm-releases.txt

Inventory checklist:

  • All PVCs using standard, standard-rwo, or premium-rwo StorageClasses
  • All ServiceAccounts with the iam.gke.io/gcp-service-account annotation, plus direct IAM principal bindings (Workload Identity Federation for GKE) that carry no annotation
  • All Config Connector CRDs (cnrm.cloud.google.com resources) and which GCP resources they manage
  • All GKE ManagedCertificate resources
  • All BackendConfig / FrontendConfig resources
  • Anthos Service Mesh configuration (VirtualService, DestinationRule)
  • Autopilot constraints: DaemonSets, privileged containers, host network
  • Artifact Registry (REGION-docker.pkg.dev) image references
  • Regional PD (multi-zone) PVCs
  • Filestore (RWX) PVCs

4. Provision Kubernetes on Quake AI#

This guide documents both provisioning paths: Magnum in section 4.1 and self-managed Kubernetes in section 4.2. Both are supported; choose based on how closely you need to match your source cluster's Kubernetes version and CNI.

4.1 Magnum#

To provision a cluster from a Magnum template, the platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API. The Kubernetes version comes from the template catalog; if you need a version or CNI it does not cover, use the self-managed path in section 4.2.

bash
openstack coe cluster template list
openstack coe cluster create production-k8s \
  --cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
  --master-count 3 \
  --node-count 3 \
  --keypair MY_KEYPAIR

For the full walkthrough including Console and quota guidance, see Create a Kubernetes cluster. To roll your own template (custom flavors, alternate Kubernetes version), see Create a cluster template.

Once the cluster reports CREATE_COMPLETE, fetch the kubeconfig:

bash
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclass

The Cinder CSI driver and the OpenStack cloud-provider stack are pre-installed by the template; no manual CCM, CSI, or CNI install is required for the Magnum path.

Build the cluster yourself on Nova instances for full control over the Kubernetes version, CNI, and CRI. This is the recommended path for migrated workloads. Use OpenTofu with the openstack provider to create the infrastructure, then install Kubernetes.

Infrastructure (OpenTofu):

  • Neutron network + subnet
  • Security groups (K8s API 6443, NodePort 30000-32767, inter-node)
  • Nova instances: 3 control plane + N workers
  • Stable endpoint for the Kubernetes API

Install RKE2 (one supported self-managed distribution):

bash
curl -sfL https://get.rke2.io | INSTALL_RKE2_TYPE=server sh -
systemctl enable --now rke2-server

Deploy the OpenStack cloud-provider stack on the self-managed cluster:

  1. openstack-cloud-controller-manager with application credentials
  2. Cinder CSI driver with cinder-flash StorageClass
  3. Calico (VXLAN) or Cilium as the CNI

5. Adapt provider-specific resources#

CSI driver: GCE persistent disk to Cinder#

Replace all StorageClass references:

YAML
# Before (GKE): standard, standard-rwo, or premium-rwo
storageClassName: standard

# After (Quake AI)
storageClassName: cinder-flash

Regional PDs: GKE supports multi-zone persistent disks for HA StatefulSets. Cinder volumes on Quake AI are single-AZ. For HA stateful workloads, implement replication at the application layer (database streaming replication, Redis Sentinel) rather than relying on multi-AZ disk replication.

Filestore (RWX): GKE Filestore provides ReadWriteMany access. Options on Quake AI:

  • NFS server provisioner (dedicated Nova VM)
  • Rook-Ceph with CephFS (production-grade RWX)
  • Application restructuring to use RWO volumes

Ingress controller: GKE ingress to nginx-ingress#

GKE Ingress uses Google's HTTP Load Balancer (GCLB) with GKE-specific CRDs. Replace with nginx-ingress:

YAML
# Before (GKE Ingress)
annotations:
  kubernetes.io/ingress.class: "gce"
  kubernetes.io/ingress.global-static-ip-name: "my-static-ip"

# After (Quake AI nginx-ingress)
annotations:
  kubernetes.io/ingress.class: nginx

Remove GKE-specific CRDs:

  • ManagedCertificate: replace with cert-manager Certificate resources using Let's Encrypt
  • BackendConfig: remove; configure health checks and timeouts via nginx-ingress annotations
  • FrontendConfig: remove; configure redirects via nginx-ingress annotations

Pod identity: workload identity to application credentials#

GKE renamed this feature to Workload Identity Federation for GKE in 2024. It has two approaches: ServiceAccount impersonation, which annotates a ServiceAccount with iam.gke.io/gcp-service-account (the older approach), and direct IAM principal binding, which grants permissions to Kubernetes resources through a principal identifier and carries no ServiceAccount annotation. Both exchange Kubernetes tokens for GCP access through Google's security token service, and both are GCP-specific and cannot work outside GKE. Audit both paths: check ServiceAccount annotations and the project IAM policy (gcloud projects get-iam-policy PROJECT_ID) for GKE principal bindings.

For applications migrating to OpenStack services (GCS to Swift, Pub/Sub to RabbitMQ): Use application credentials in K8s Secrets.

For applications that still need GCP access (hybrid period): Use service account JSON keys:

bash
kubectl create secret generic gcp-sa-key \
  --from-file=key.json=/path/to/service-account.json

Mount as a volume and set GOOGLE_APPLICATION_CREDENTIALS to the mount path.

Remove the Workload Identity annotation from every affected ServiceAccount:

bash
kubectl annotate sa SA_NAME iam.gke.io/gcp-service-account- -n NAMESPACE

Config connector cRDs#

Config Connector creates GCP resources (BigQuery datasets, Cloud SQL instances, Pub/Sub topics) from K8s manifests. These CRDs only function in GKE clusters with GCP access.

Remove all Config Connector resources before restoring on Quake AI:

bash
kubectl get crds | grep cnrm.cloud.google.com | awk '{print $1}' | xargs kubectl delete crd

Provision equivalent OpenStack resources via OpenTofu instead.

Service type LoadBalancer#

Remove GKE-specific annotations. The cloud controller assigns a public endpoint to a Kubernetes Service with type: LoadBalancer.

Autopilot constraints#

If migrating from GKE Autopilot:

  • Deploy the CCM, CSI, and CNI DaemonSets that Autopilot blocked from kube-system. Autopilot restricts DaemonSets only in GKE-managed namespaces, not user-deployed DaemonSets, so these run normally on Quake AI.
  • Allow privileged containers where needed
  • Remove node.kubernetes.io/workload tolerations that were Autopilot-specific
  • Remove minimum resource request constraints imposed by Autopilot

Container images: Artifact Registry#

If using Artifact Registry (REGION-docker.pkg.dev) with GKE-integrated authentication, mirror images to a permanent registry (Docker Hub, Harbor). Artifact Registry auth tokens tied to GKE metadata do not work from Quake AI. Google Container Registry (gcr.io) was shut down in March 2025; any remaining images should already have moved to Artifact Registry.

Anthos Service mesh#

If using Anthos Service Mesh (ASM), uninstall it and install upstream Istio or Linkerd. ASM CRDs differ from upstream Istio. Re-apply service mesh policies (VirtualService, DestinationRule) using upstream Istio CRDs.

6. Apply and validate#

Migrate data with Velero:

bash
# On GKE: install Velero with GCP backend
velero install \
  --provider gcp \
  --plugins velero/velero-plugin-for-gcp:v1.9.0 \
  --bucket gke-migration-backup \
  --use-node-agent \
  --default-volumes-to-fs-backup

velero backup create gke-full --include-namespaces production

# On Quake AI: restore with StorageClass remapping
kubectl apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/plugin-config: ""
    velero.io/change-storage-class: RestoreItemAction
data:
  standard: cinder-flash
  standard-rwo: cinder-flash
  premium-rwo: cinder-flash
EOF

velero restore create --from-backup gke-full --restore-volumes=true

Deploy and verify:

bash
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces
kubectl get ingress --all-namespaces

For Cloud SQL databases, use managed export/import to self-hosted PostgreSQL or MySQL on Quake AI. For GCS data, use rclone copy gs://bucket s3://swift-bucket.

7. Observability setup#

Replace Google Cloud Operations (Kubernetes Engine Monitoring, Cloud Logging, Cloud Trace) with a self-managed stack:

bash
helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set promtail.enabled=true \
  --set loki.persistence.storageClassName=cinder-flash

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the kube-prometheus-stack chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin --version 72.9.1, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the --version flag to install the current chart.

For distributed tracing (replacing Cloud Trace), add Grafana Tempo.

8. Auto-scaling alternatives#

GKE featureQuake AI equivalent
Autopilot (fully managed nodes, no node pools)No equivalent. Pre-provision node pools in OpenTofu and configure the Cluster Autoscaler. Size nodes from your workload resource requests, which Autopilot sized automatically
Cluster Autoscaler (Standard mode)Cluster Autoscaler with OpenStack provider
Node Auto-Provisioning (NAP)Not available; pre-provision node pools of different flavors
VPAFully portable
HPAFully portable

If migrating from Autopilot, you must explicitly size and manage worker nodes. Pre-provision node pools with OpenTofu and configure the Cluster Autoscaler with --cloud-provider=openstack.

9. Validation checklist#

  • All pods in Running or Completed state
  • PVCs bound to Cinder volumes with correct data
  • Ingress routes working through the ingress controller's public LoadBalancer Service
  • DNS resolves to the new Service address
  • cert-manager certificates issued (replacing ManagedCertificate)
  • No iam.gke.io annotations remaining on ServiceAccounts
  • No Config Connector CRDs (cnrm.cloud.google.com) remaining
  • No BackendConfig/FrontendConfig/ManagedCertificate CRDs remaining
  • Prometheus scraping all targets
  • Loki ingesting logs
  • StatefulSet data integrity verified
  • Application health checks passing
  • DaemonSets running (if migrated from Autopilot)

10. Provider-specific gotchas#

GotchaImpactMitigation
Autopilot restricts DaemonSets in kube-systemCCM, CSI, and CNI must run as DaemonSetsDeploy them yourself on Quake AI; user DaemonSets run normally
Workload Identity Federation is GCP-onlyOIDC token exchange doesn't work outside GCPReplace with service account JSON keys or migrate to OpenStack services
Config Connector CRDs only work in GKECRDs become invalid on non-GKE clustersRemove before restore; provision equivalents via OpenTofu
GKE Backup (Backup for GKE) format is GCP-nativeCannot restore to Quake AI directlyUse Velero + filesystem backup instead
Regional PDs (multi-zone HA storage)Cinder is single-AZ per volumeImplement application-level HA (database replication)
GKE Ingress creates Cloud Armor WAF rulesThe destination cluster needs an application-layer WAFDeploy ModSecurity with ingress-nginx, or place an external CDN and WAF in front of the Service address
cloud.google.com/ node labelsMay break nodeAffinity scheduling rulesUpdate nodeAffinity to use standard K8s labels
Anthos Service Mesh CRDs differ from upstream IstioASM VirtualService/DestinationRule may need adjustmentUninstall ASM; install upstream Istio; re-apply policies
GKE Enterprise Config Management (GitOps) has no Quake AI equivalentPolicy-as-code and config sync stop working off GKEReplace with Argo CD or Flux
GKE Enterprise Policy Controller (OPA-based) has no Quake AI equivalentAdmission policies stop enforcing off GKEReplace with OPA Gatekeeper or Kyverno (both install via Helm)
Filestore (RWX) has no direct Cinder equivalentApplications requiring shared storage need restructuringNFS server provisioner, Rook-Ceph, or application-level coordination

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

Comparisons to third-party providers in this material reflect publicly documented behavior as of the validation date below. Pricing, quotas, service limits, and feature availability change frequently on every cloud. Verify provider-specific claims against the provider's own current documentation before relying on them for a procurement, architecture, or migration decision.

For the full policy, see Usage Guidelines.

Last validated: 22.06.2026

Before this

Quick answers

Was this page helpful?