Skip to content

Migrate from Vultr Kubernetes Engine (VKE) to Kubernetes on Quake AI

Migration

Coming from another cloud?

▸Vultr·Kubernetes Engine (VKE) Cluster

Vultr Kubernetes Engine (VKE) Clusterhigh

  • VKE provides a free managed control plane with optional HA upgrades; on Quake AI you provision the cluster through Magnum or self-managed Kubernetes on Nova instances (OpenTofu plus kubeadm, k3s, or RKE2) and operate it yourself.
  • VKE clusters are created and managed through the Vultr API or Customer Portal with a single cluster object. Quake AI offers Magnum as a managed-cluster API, or you can compose a cluster from compute instances, a private network, a CNI, and a self-managed control-plane endpoint.
  • VKE ships an integrated Vultr cloud controller manager and Vultr CSI driver for Block Storage out of the box; on Quake AI you install openstack-cloud-controller-manager and the Cinder CSI driver yourself.
  • VKE node pools auto-recycle when you upgrade the Kubernetes version; Quake AI node lifecycle is managed by you (OpenTofu replace, manual drain, or cluster API operator).
Vultr docs ↗

Migrate from Vultr Kubernetes Engine (VKE) to Kubernetes on Quake AI

Service mapping#

Vultr to Quake AI

Vultr serviceQuake AI equivalentKey difference
BucketsContainersVultr calls the OpenStack Swift container equivalent a bucket, and the docs describe buckets as the primary organizational unit for storing... see details
Kubernetes Engine (VKE) ClusterKubernetesVKE provides a free managed control plane with optional HA upgrades; on Quake AI you provision the cluster through Magnum or self-managed... see details

1. Overview#

Vultr Kubernetes Engine (VKE) is Vultr's managed Kubernetes service. VKE ships vanilla upstream Kubernetes with a small set of Vultr-specific add-ons: a Vultr cloud controller manager (handles Service type: LoadBalancer provisioning of Vultr Load Balancers and node lifecycle), a Vultr Block Storage CSI driver, and integration with the Vultr Container Registry for images. VKE offers a free control plane (single region) and optional HA control plane upgrades.

On Quake AI, you provision Kubernetes through Magnum, the Container Infrastructure Management service. Magnum gives you a control plane and worker nodes from a curated cluster template (openstack coe cluster create). The us-east-1 region ships platform templates pre-wired with Calico CNI, the OpenStack cloud-provider stack (CCM + Cinder CSI), and a load-balanced API endpoint. Run openstack coe cluster template list for the current template catalog and Kubernetes versions, or see the cluster templates reference.

For teams that need a non-Magnum stack (custom CNI, custom kubelet flags, custom CRI, or a Kubernetes version outside the template catalog), self-managed RKE2 or k3s on Nova instances is documented as an alternative later in this tutorial. VKE users accustomed to a fully managed control plane should plan for the operational shift: on Magnum you own the cluster after creation (upgrades, node-pool sizing, etcd backups), and on the self-managed path you also own the bootstrap, the CNI, and the CCM/CSI install.

Migration complexity: Low to medium. VKE has no IRSA-style pod identity, no proprietary CNI, and only a handful of vendor-specific annotations. The primary tasks are swapping the CCM and CSI driver, mirroring container images, and rebuilding a stable Kubernetes API endpoint. Expect 1 to 2 weeks for a typical production migration.

2. Workload portability matrix#

Resource typePortabilityNotes
DeploymentPortableNo changes needed
StatefulSet (manifest)PortableData migrated separately
DaemonSetPortableNo changes needed
ConfigMapPortableUpdate any Vultr-specific endpoint values
Secret (values)PortableRemove VULTR_API_KEY references
Service (ClusterIP, NodePort)PortableNo changes
Service (LoadBalancer)Needs adaptationRemove service.beta.kubernetes.io/vultr-loadbalancer-* annotations
IngressPortablenginx-ingress is commonly used on VKE already
PersistentVolumeClaimNeeds adaptationChange storageClassName from vultr-block-storage to cinder-flash
PersistentVolume (data)Not portableMigrate data via Velero filesystem backup or application-level dump/restore
NetworkPolicyPortableVKE supports standard NetworkPolicy through Calico/Cilium-compatible CNIs
HPA / VPA / PDBPortableNo changes
CronJob / JobPortableNo changes
RBAC resourcesPortableNo changes
Vultr CSI driver (block.csi.vultr.com)Provider-specificReplace with Cinder CSI
Vultr CCM (vultr)Provider-specificReplace with openstack-cloud-controller-manager
Vultr Container Registry imagesProvider-specificMirror to Docker Hub, Harbor, or Quay.io

3. Pre-migration: export and audit#

bash
kubectl get all --all-namespaces -o yaml > cluster-export.yaml

kubectl get pvc,pv --all-namespaces -o yaml > storage.yaml

kubectl get svc --all-namespaces -o yaml | grep 'vultr-loadbalancer' > vultr-lb-services.txt

helm list --all-namespaces > helm-releases.txt

Inventory checklist:

  • All PVCs using vultr-block-storage StorageClass
  • All Services with service.beta.kubernetes.io/vultr-loadbalancer-* annotations
  • Images hosted on the Vultr Container Registry
  • Helm releases and versions
  • external-dns configuration (if using a managed DNS provider)
  • Any references to VULTR_API_KEY in Secrets

4. Provision Kubernetes on Quake AI#

4.1 Primary path: Magnum#

Provision a cluster from a Magnum template. The platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API.

bash
openstack coe cluster template list
openstack coe cluster create production-k8s \
  --cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
  --master-count 3 \
  --node-count 3 \
  --keypair MY_KEYPAIR

For the full walkthrough including Console and quota guidance, see Create a Kubernetes cluster. To roll your own template (custom flavors, alternate Kubernetes version), see Create a cluster template.

Once the cluster reports CREATE_COMPLETE, fetch the kubeconfig:

bash
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclass

The template pre-installs the Cinder CSI driver and the OpenStack cloud-provider stack; the Magnum path needs no manual CCM, CSI, or CNI install.

4.2 Alternative: self-managed RKE2 or k3s#

If the Magnum template catalog does not match your version, CNI, or CRI constraints, build the cluster yourself on Nova instances. Use OpenTofu with the openstack provider to create the infrastructure:

  • Neutron network + subnet
  • Security groups (Kubernetes API 6443, NodePort 30000-32767, inter-node)
  • Nova instances: 3 control plane + N workers
  • Stable endpoint for the Kubernetes API

Install RKE2 or k3s, then deploy on the cluster:

  1. openstack-cloud-controller-manager with application credentials
  2. Cinder CSI driver with the cinder-flash StorageClass
  3. Calico (VXLAN) or Cilium as the CNI
bash
# Example: minimal RKE2 control plane bootstrap on a Nova instance
curl -sfL https://get.rke2.io | sh -
sudo systemctl enable --now rke2-server
sudo cat /etc/rancher/rke2/rke2.yaml  # kubeconfig

Point the kubeconfig server: value at the control-plane endpoint on port 6443.

5. Adapt provider-specific resources#

CSI driver: Vultr Block Storage to Cinder#

Replace the StorageClass reference in all PVCs:

YAML
# Before (VKE)
storageClassName: vultr-block-storage

# After (Quake AI)
storageClassName: cinder-flash

VKE volumes are standard RWO block storage; Cinder is a 1:1 replacement. If you depended on a retain-on-delete behavior, create or use a Cinder StorageClass with reclaimPolicy: Retain.

Ingress controller#

If your VKE cluster already uses nginx-ingress (the most common pattern), no ingress changes are needed beyond DNS updates. nginx-ingress is fully portable.

If you used the Vultr cloud controller manager to provision Vultr Load Balancers directly through Service type: LoadBalancer, remove Vultr-specific annotations:

YAML
# Remove these Vultr-specific annotations from Service manifests
service.beta.kubernetes.io/vultr-loadbalancer-algorithm
service.beta.kubernetes.io/vultr-loadbalancer-protocol
service.beta.kubernetes.io/vultr-loadbalancer-ssl
service.beta.kubernetes.io/vultr-loadbalancer-sticky-sessions
service.beta.kubernetes.io/vultr-loadbalancer-healthcheck-path
service.beta.kubernetes.io/vultr-loadbalancer-firewall-rules

The cloud controller assigns public endpoints to Kubernetes Services with type: LoadBalancer without requiring provider-specific annotations.

Pod identity#

VKE has no IAM-for-pods mechanism. Pods on VKE that call Vultr APIs (for example, external-dns with the Vultr DNS provider) use long-lived API keys stored in K8s Secrets.

  • Remove any VULTR_API_KEY Secret references that are not still required.
  • For applications that called Vultr APIs, reconfigure to point at the target service:
YAML
# Before: external-dns with Vultr DNS
args:
  - --provider=vultr

# After: external-dns with Cloudflare (or other)
args:
  - --provider=cloudflare

Service type LoadBalancer#

The standard Kubernetes LoadBalancer Service fields remain portable. Remove any vultr-loadbalancer-* annotations before applying on Quake AI.

Container images: Vultr Container Registry to a portable registry#

The Vultr Container Registry is Vultr-specific. Mirror all images before cutover:

bash
docker pull <region>.vultrcr.com/MY_NAMESPACE/MY_IMAGE:TAG
docker tag  <region>.vultrcr.com/MY_NAMESPACE/MY_IMAGE:TAG docker.io/ORG/MY_IMAGE:TAG
docker push docker.io/ORG/MY_IMAGE:TAG

Update all Deployment and StatefulSet manifests with the new image references. This is the most manual step in a VKE migration.

6. Apply and validate#

Migrate data with Velero:

bash
# On VKE: install Velero with an S3-compatible backend
velero install \
  --provider aws \
  --plugins velero/velero-plugin-for-aws:v1.9.0 \
  --bucket vke-migration \
  --use-node-agent \
  --default-volumes-to-fs-backup \
  --backup-location-config \
    region=ewr,s3ForcePathStyle=true,s3Url=https://ewr1.vultrobjects.com

velero backup create vke-full --include-namespaces production

# On Quake AI: restore with StorageClass remapping
kubectl apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/plugin-config: ""
    velero.io/change-storage-class: RestoreItemAction
data:
  vultr-block-storage: cinder-flash
EOF

velero restore create --from-backup vke-full --restore-volumes=true

Deploy and verify:

bash
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces

For stateful workloads (databases), prefer application-level dump/restore (pg_dump, mysqldump) over Velero filesystem backup for guaranteed consistency.

7. Observability setup#

Most VKE users already run self-managed Prometheus + Grafana. If so, these migrate as-is; only PVC StorageClass references need updating.

If starting fresh, install the self-managed observability stack:

bash
helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set promtail.enabled=true \
  --set loki.persistence.storageClassName=cinder-flash

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the kube-prometheus-stack chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin --version 72.9.1, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the --version flag to install the current chart.

VKE's basic node and cluster metrics in the Vultr Customer Portal have no equivalent on Quake AI. All monitoring is self-managed.

8. Auto-scaling alternatives#

VKE featureQuake AI equivalent
Managed node-pool autoscaling (VKE API)Cluster Autoscaler with OpenStack provider, or manual Nova scaling via OpenTofu
HPAFully portable
VPAFully portable

VKE node-pool autoscaling supports min/max bounds per pool. The equivalent on Quake AI is the OpenStack Cluster Autoscaler or manual scaling of the OpenTofu-managed worker pool.

9. Validation checklist#

  • All pods in Running or Completed state
  • PVCs bound to Cinder volumes with correct data
  • Ingress routes working through the ingress controller's public LoadBalancer Service
  • DNS resolves to new Floating IPs
  • Prometheus scraping all targets
  • No vultr-loadbalancer-* annotations remaining
  • No Vultr Container Registry image references remaining
  • No VULTR_API_KEY Secrets remaining (unless an external Vultr integration is still required)
  • StatefulSet data integrity verified
  • Application health checks passing

10. Provider-specific gotchas#

GotchaImpactMitigation
Vultr Container Registry credentials are account-scopedQuake AI cannot pull from VCR without active credentialsMirror all images before migration
Vultr-specific LB annotations are ignored by the OpenStack CCMLB may not have the desired configurationReview and remove vultr-loadbalancer-* annotations
VKE-injected operator labels / annotationsVultr-injected metadata may persist in exportsStrip Vultr-injected metadata before restoring on Quake AI
VKE HA control plane SLAQuake AI has no managed-control-plane SLAPlan control-plane HA with 3 control-plane Nova instances across host aggregates, a stable control-plane endpoint, and regular etcd backups
VKE auto-upgradesQuake AI does not auto-upgrade KubernetesTrack upstream releases and schedule kubeadm/k3s/RKE2 upgrades on your own cadence
Vultr Load Balancer health-check semanticsThe destination ingress controller may behave differentlyRe-test health checks under load after migration

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

Comparisons to third-party providers in this material reflect publicly documented behavior as of the validation date below. Pricing, quotas, service limits, and feature availability change frequently on every cloud. Verify provider-specific claims against the provider's own current documentation before relying on them for a procurement, architecture, or migration decision.

For the full policy, see Usage Guidelines.

Before this

Quick answers

Was this page helpful?