Migrate a Kubernetes platform to Quake AI
Coming from another cloud?
▸AWS·Amazon EKS Cluster
Amazon EKS Cluster
- EKS control plane fully AWS-managed, single-tenant, across 3 AZs with auto scale/replace; on Quake AI you run the control plane on Nova instances, provisioned through Magnum or self-managed with OpenTofu, and you operate it.
- EKS regional API endpoint with SLA; Quake AI exposes the kube API via Neutron LB with floating IP.
- EKS charges a per-hour cluster platform fee on top of the underlying compute; Quake AI charges only for underlying Nova/Neutron/Cinder resources with no K8s platform fee.
- EKS managed nodes auto AMI updates, Spot integration; Quake AI self-managed nodes require manual OS image selection and update management.
▸Azure·AKS Cluster
AKS Cluster
- Azure automatically provisions and manages the control plane at no additional cost (Free tier) or fixed fee (Standard tier with SLA), offloading health monitoring and upgrades; on Quake AI you provision a cluster through Magnum (openstack coe cluster create) or self-managed Kubernetes on Nova instances (OpenTofu plus kubeadm, k3s, or RKE2), and you operate the cluster after creation.
- No OpenStack integration; uses Azure Resource Manager for cluster lifecycle.
- Pre-configured with Azure-specific defaults and add-ons like application routing.
- Managed via Azure Virtual Machine Scale Sets (VMSS) with auto-scaling and upgrades; Quake AI uses Nova instances provisioned via OpenTofu with user-managed scaling and upgrades.
▸DigitalOcean·K8S
This Quake AI feature maps to DigitalOcean’s K8S.
▸Google Cloud·GKE Cluster
GKE Cluster
- GKE provides Autopilot mode with fully managed node provisioning and scaling by Google; on Quake AI you provision clusters through Magnum or self-managed Kubernetes on Nova instances and manage node scaling yourself.
- Control plane is fully managed with automatic upgrades through release channels; Quake AI requires user-provisioned and user-managed control plane nodes.
- Cluster creation uses gcloud CLI vs OpenStack CLI (openstack coe cluster create).
- Custom machine types, spot VMs, accelerators in node pools; Quake AI K8s nodes use standard Nova flavors.
Migrate a Kubernetes platform to Quake AI
In this migration guide, you move an existing Kubernetes platform to Quake AI. You stand up the Kubernetes cluster template on Magnum, sync container images and persistent data from the source cluster, and cut over ingress and DNS with a workload drain plan and rollback.
Quake AI runs a managed control plane through Magnum; worker nodes are Nova instances you size and patch. You operate cluster add-ons, ingress, and Helm releases on top of Magnum.
What you will migrate: a platform on Amazon EKS, Google GKE, Azure AKS, DigitalOcean Kubernetes (DOKS), or a self-managed cluster elsewhere.
What you will learn:
- How to provision the
k8s-clustertemplate as the migration target - How to pick the correct primitive migration page (EKS, GKE, AKS, or DOKS)
- How to mirror images and sync backup objects with Migrate from S3
- How to cut over ingress, drain source nodes, and roll back if workloads fail on Magnum
Time estimate: 75 minutes (excluding image mirror and PV copy time)
Prerequisites#
Before you start, confirm you have:
kubectlaccess to the source cluster and cluster-admin on the target Magnum cluster after provision- OpenTofu installed locally and application credentials for Quake AI
- Enough quota for a Magnum cluster, worker nodes, and the public IPs your ingress design requires
- The concept-translation page for your source: Coming from AWS, Coming from Azure, Coming from GCP, or Coming from DigitalOcean
Skim Deploy the Kubernetes cluster template with OpenTofu if you have not provisioned Magnum through the template before.
Step 1: Map the source provider#
Inventory namespaces, ingress hostnames, storage classes, container registries, and StatefulSets with PersistentVolumeClaims. Note API version differences between the source distribution and Magnum.
Step 2: Stand up the target cluster#
- Download k8s-cluster.zip from the Kubernetes Cluster Bootstrap template page and unzip it into a working directory.
- Set cluster name, node count, and flavor variables in
terraform.tfvars. - Run
tofu applyand configurekubectlagainst the Magnum kubeconfig output.
Record how many floating IPs your ingress design allocates. The template reference page documents the default count; add public addresses for LoadBalancer services or edge-proxy VMs when your architecture requires them.
Step 3: Move Kubernetes workloads#
Follow the primitive migration page for your source control plane:
| Source | Migration page |
|---|---|
| Amazon EKS | Migrate from EKS |
| Google GKE | Migrate from GKE |
| Azure AKS | Migrate from AKS |
| DigitalOcean DOKS | Migrate from DOKS |
Those pages cover manifest export, API mapping, and workload rescheduling. Run helm template or kubectl diff against Magnum before you drain production traffic.
Step 4: Move images and artifact data#
Mirror container images to a registry Magnum worker nodes can reach (Quake AI-hosted registry, Docker Hub, or a registry you operate). Update image references in manifests before cutover.
Sync Helm chart archives, backup tarballs, and checkpoint objects with Migrate from S3 when your pipeline stores artifacts in object storage.
Step 5: Migrate persistent volume data#
Snapshot or copy data bound to PersistentVolumeClaims before you reschedule stateful workloads. Validate restore on Magnum worker nodes with matching storage classes. Update PVC manifests to reference Magnum-compatible classes.
Step 6: Cut over ingress and DNS#
- Deploy or reconfigure your ingress controller on Magnum. Publish it through a Kubernetes
LoadBalancerservice or an edge proxy on a floating IP. - Configure TLS certificates through your ingress controller or certificate operator.
- Lower DNS TTL in advance, then point hostnames at the Magnum ingress address when health checks pass.
Workload drain and rollback: cordon and drain source nodes only after Magnum passes health checks and error budgets hold. If error rates spike, switch DNS back to the source ingress and uncordon source nodes.
Step 7: Verify the migrated platform#
- Confirm critical Deployments reach
Readystate on Magnum. - Run application smoke tests against production hostnames.
- Validate PVC mounts and database connectivity for stateful services.
Step 8: Clean up migration scaffolding#
Delete temporary mirror registries, drop staging namespaces, and decommission the source cluster only after traffic stabilizes on Magnum for the window your team requires.
What you migrated#
You moved a Kubernetes platform to Quake AI Magnum with the Kubernetes cluster template as the target, composed primitive Kubernetes and object migration pages, and cut over ingress with an explicit drain and rollback plan.
Return to the Kubernetes platforms solutions leaf for the workload-shaped migration overview.
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
Comparisons to third-party providers in this material reflect publicly documented behavior as of the validation date below. Pricing, quotas, service limits, and feature availability change frequently on every cloud. Verify provider-specific claims against the provider's own current documentation before relying on them for a procurement, architecture, or migration decision.
For the full policy, see Usage Guidelines.
Last validated: 08.09.2026
Quick answers
- Why does `openstack coe cluster create` fail with a Keystone trust or unauthorized error when I use an application credential?CLIAPITerraform
- Why does a Kubernetes LoadBalancer service stay `<pending>` for several minutes?CLI
- Why does my GitHub Actions or GitLab CI job fail to run `openstack coe` or `kubectl` on a Magnum cluster?CLI
- Why does my Magnum cluster create fail with "Only volume-backed servers" or "Quota exceeded for compute_units"?CLIAPI