# Migrate from Azure AKS to Kubernetes on Quake AI

Source: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-aks
Markdown: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-aks.md

---

# Migrate from Azure AKS to Kubernetes on Quake AI

## Service mapping

<MigrationTable provider="azure" service="kubernetes" />


## 1. Overview

Azure Kubernetes Service (AKS) provides a managed Kubernetes control plane with deep Azure integration: a choice of container network interface (Azure CNI with VNET IPs, Azure CNI Overlay with a separate pod CIDR and NAT for external traffic, or Azure CNI with the Cilium dataplane), Microsoft Entra Workload Identity and the deprecated AAD Pod Identity for pod-level IAM, Azure Key Vault CSI for secrets, Application Gateway Ingress Controller (AGIC) or Application Gateway for Containers (AGfC) for L7 ingress, Azure Managed Disks and Azure Files for storage, and Azure Monitor Container Insights for observability. AKS Automatic clusters use Azure CNI Overlay with the Cilium dataplane by default.

On Quake AI, you run Kubernetes on Nova instances through one of two supported paths. You can provision a cluster through [Magnum](/docs/kubernetes), the Container Infrastructure Management service, which builds a control plane and worker nodes from a curated cluster template (`openstack coe cluster create`), with the Kubernetes version set by the template. The platform-provided `Standard-V2.0-k8s-calico-fc38_*` templates cover the currently supported versions and ship pre-wired with Calico CNI, the OpenStack cloud-provider stack, and a load-balanced API endpoint. Or you can run **self-managed RKE2 or k3s** with the OpenStack cloud-provider stack (CCM + Cinder CSI), which keeps the Kubernetes version, CNI, and CRI under your control. Run `openstack coe cluster template list` to see the versions your region offers rather than assuming a specific version from this guide.

AKS offers Long-Term Support (LTS) for selected Kubernetes versions (1.29 onward), so some teams run older versions deliberately. Confirm your target Kubernetes version against the available Magnum templates, or take the **self-managed RKE2 or k3s on Nova instances** path for a version outside the template catalog or for custom CNI, kubelet, or CRI requirements. Both paths are covered later in this tutorial.

**Migration complexity: High.** This is the most complex migration of the five providers. Azure CNI is tightly coupled to VNET, Microsoft Entra Workload Identity must be replaced with OpenStack application credentials, Azure Key Vault secrets must be pre-extracted, and the AGIC-to-nginx-ingress conversion is the most annotation-heavy of the five. Expect **6-12 weeks** for large production clusters.

## 2. Workload portability matrix

| Resource type | Portability | Notes |
|---|---|---|
| Deployment | Portable | No changes needed |
| StatefulSet (manifest) | Portable | Data migrated separately via Velero |
| DaemonSet | Portable | No changes needed |
| ConfigMap | Portable | Update any Azure endpoint values |
| Secret (values) | Portable | Must re-create if sourced from Key Vault |
| Service (ClusterIP, NodePort) | Portable | No changes |
| Service (LoadBalancer) | Needs adaptation | Remove Azure-specific annotations; the cloud controller assigns the public endpoint |
| Ingress | Needs adaptation | Replace AGIC annotations with nginx-ingress |
| PersistentVolumeClaim | Needs adaptation | Change `storageClassName` from `managed-csi` to `cinder-flash` |
| PersistentVolume (data) | Not portable | Migrate via Velero filesystem backup |
| NetworkPolicy | Portable | Requires Calico/Cilium on Quake AI |
| HPA / VPA / PDB | Portable | No changes |
| CronJob / Job | Portable | No changes |
| ServiceAccount | Needs adaptation | Remove `azure.workload.identity/client-id` annotation |
| RBAC resources | Portable | No changes |
| Azure CNI / Azure CNI Overlay | Provider-specific | Replace with Calico or Cilium |
| AGIC / Application Gateway for Containers (AGfC) | Provider-specific | Replace with ingress-nginx or Traefik and a `LoadBalancer` Service |
| Azure Managed Disk CSI (`disk.csi.azure.com`) | Provider-specific | Replace with Cinder CSI |
| Azure Files CSI (`file.csi.azure.com`) | Provider-specific | Replace with NFS provisioner or Rook-Ceph |
| Azure Key Vault CSI | Provider-specific | Export secrets; use K8s Secrets or HashiCorp Vault |
| Microsoft Entra Workload Identity | Provider-specific | Remove; use OpenStack app credentials |
| AAD Pod Identity (deprecated) | Provider-specific | Remove NMI DaemonSet and CRDs |
| `omsagent`/`ama-logs` DaemonSet | Provider-specific | Remove; replace with Promtail/Fluent Bit |
| `azure-policy` pod | Provider-specific | Replace with OPA Gatekeeper or Kyverno |
| AKS Node Auto-Provisioning (NAP) | Provider-specific | Cluster Autoscaler with OpenStack provider; remove `AKSNodeClass` and `NodePool` CRDs |

## 3. Pre-migration: export and audit

```bash
kubectl get all --all-namespaces -o yaml > cluster-export.yaml

kubectl get pvc --all-namespaces -o yaml > pvcs.yaml

kubectl get sa --all-namespaces -o yaml | grep -l 'azure.workload.identity' > workload-id-accounts.txt

kubectl get azureidentity,azureidentitybinding --all-namespaces -o yaml > pod-identity.yaml

kubectl get secretproviderclass --all-namespaces -o yaml > key-vault-spc.yaml

kubectl get ingress --all-namespaces -o yaml > ingresses.yaml

helm list --all-namespaces > helm-releases.txt
```

**Inventory checklist:**

- [ ] All ServiceAccounts with `azure.workload.identity/client-id` (Workload ID)
- [ ] All AzureIdentity / AzureIdentityBinding resources (legacy Pod Identity)
- [ ] All SecretProviderClass resources (Azure Key Vault CSI)
- [ ] All PVCs using `managed-csi`, `managed-csi-premium`, `azurefile-csi`
- [ ] All AGIC Ingress resources and their Application Gateway annotations
- [ ] Azure Files (SMB/NFS) PVCs: plan replacement
- [ ] Azure DevOps pipeline connections to this cluster
- [ ] `omsagent` / `ama-logs` DaemonSets
- [ ] ACR (Azure Container Registry) image references

## 4. Provision Kubernetes on Quake AI

### 4.1 Magnum

This section covers the Magnum path. If you need a Kubernetes version, CNI, or CRI outside the template catalog, use the self-managed RKE2 or k3s path in section 4.2 instead.

Provision a cluster from a Magnum template. The platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API.


**Magnum cluster creation requires a password-scoped session.** `openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Trust delegation is not available to Keystone application credentials, so the call fails before any Magnum work begins. Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` (password auth) before running the command. The Cloud Console wizard works from any logged-in user session because it uses the user's password-scoped token. See [the Kubernetes FAQ](/docs/kubernetes/faq) for the full troubleshooting flow.


```bash
openstack coe cluster template list
openstack coe cluster create production-k8s \
  --cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
  --master-count 3 \
  --node-count 3 \
  --keypair MY_KEYPAIR
```

For the full walkthrough including Console and quota guidance, see [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster). To roll your own template (custom flavors, alternate Kubernetes version), see [Create a cluster template](/docs/kubernetes/how-to/create-cluster-template).

Once the cluster reports `CREATE_COMPLETE`, fetch the kubeconfig:

```bash
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclass
```

The Cinder CSI driver and the OpenStack cloud-provider stack are pre-installed by the template; no manual CCM, CSI, or CNI install is required for the Magnum path.

### 4.2 Recommended: self-managed RKE2 or k3s

This is the recommended path for new and migrated workloads, and the only option when the Magnum template catalog does not match your version, CNI, or CRI constraints. Build the cluster yourself on Nova instances using OpenTofu with the `openstack` provider to create the infrastructure:

- Neutron network + subnet
- Security groups (K8s API 6443, NodePort 30000-32767, inter-node)
- Nova instances: 3 control plane + N workers
- Stable endpoint for the Kubernetes API

**Install RKE2 (one supported self-managed distribution):**

```bash
curl -sfL https://get.rke2.io | INSTALL_RKE2_TYPE=server sh -
systemctl enable --now rke2-server
```

Deploy the OpenStack cloud-provider stack on the self-managed cluster:
1. `openstack-cloud-controller-manager` with [application credentials](/docs/identity/how-to/create-application-credential)
2. Cinder CSI driver with `cinder-flash` StorageClass
3. Calico (VXLAN) or Cilium as the CNI

## 5. Adapt provider-specific resources

### Prerequisite: export Azure Key Vault secrets

This step must complete before application pods can start on Quake AI. Export all secrets from Azure Key Vault and create K8s Secrets:

```bash
az keyvault secret list --vault-name my-vault --query "[].name" -o tsv | \
  while read secret; do
    value=$(az keyvault secret show --vault-name my-vault --name "$secret" --query value -o tsv)
    kubectl create secret generic "kv-$secret" --from-literal=value="$value"
  done
```

Remove all `SecretProviderClass` resources and Key Vault CSI volume references from pod specs. For enterprise secret management, deploy HashiCorp Vault on Quake AI instead of K8s Secrets.

### CSI driver: Azure managed disks to Cinder

Replace all StorageClass references:

```yaml
# Before (AKS): managed-csi or managed-csi-premium
storageClassName: managed-csi

# After (Quake AI)
storageClassName: cinder-flash
```

**Azure Files (RWX):** Azure Files provides ReadWriteMany via NFS or SMB. Options on Quake AI:
- NFS server on a dedicated Nova VM
- Rook-Ceph with CephFS (production-grade but operationally complex)
- For Azure Files data, use `rsync` to an NFS server on Quake AI

### Ingress controller: AGIC to nginx-ingress

Replace Application Gateway Ingress Controller annotations:

```yaml
# Before (AKS AGIC)
annotations:
  kubernetes.io/ingress.class: azure/application-gateway
  appgw.ingress.kubernetes.io/ssl-redirect: "true"
  appgw.ingress.kubernetes.io/backend-path-prefix: "/"

# After (Quake AI nginx-ingress)
annotations:
  kubernetes.io/ingress.class: nginx
  nginx.ingress.kubernetes.io/ssl-redirect: "true"
```

Deploy cert-manager with a Let's Encrypt ClusterIssuer to replace Azure Application Gateway's managed TLS termination.

If your source cluster uses Application Gateway for Containers (AGfC, newer than AGIC), the migration target is the same: an in-cluster ingress controller exposed through a `LoadBalancer` Service. AGfC uses different CRDs and Gateway API resources, so your source manifests will differ from the AGIC annotations shown above.

### Pod identity: Entra ID to application credentials

AKS has two pod identity mechanisms; both are Azure-specific.

**Microsoft Entra Workload Identity (current):** Annotates ServiceAccounts with `azure.workload.identity/client-id`. Remove the annotation and the injected env vars (`AZURE_CLIENT_ID`, `AZURE_TENANT_ID`, `AZURE_FEDERATED_TOKEN_FILE`):

```bash
kubectl annotate sa SA_NAME azure.workload.identity/client-id- -n NAMESPACE
```

**AAD Pod Identity (deprecated):** Uses `AzureIdentity`, `AzureIdentityBinding`, and `AzurePodIdentityException` CRDs with an NMI DaemonSet that intercepts IMDS calls. Remove all these CRDs and the NMI DaemonSet. They will fail on non-AKS clusters.

**For applications accessing OpenStack APIs:** Use [application credentials](/docs/identity/how-to/create-application-credential) in K8s Secrets.

**For applications that still need Azure access (hybrid period):** Create a service principal with client credentials:

```bash
kubectl create secret generic azure-sp-credentials \
  --from-literal=AZURE_CLIENT_ID=CLIENT_ID \
  --from-literal=AZURE_CLIENT_SECRET=CLIENT_SECRET \
  --from-literal=AZURE_TENANT_ID=TENANT_ID
```

### CNI replacement: Azure CNI to Calico/Cilium

Azure CNI assigns pods Azure VNET IPs, while Azure CNI Overlay assigns pods IPs from a separate overlay CIDR and uses NAT for external traffic. Either way the pod network does not move: the Quake AI cluster starts fresh with its own CNI. For Overlay clusters this is a smaller conceptual change, because the pod IP space is already overlay-based, similar to Calico VXLAN.

Check the [Calico installation documentation](https://docs.tigera.io/calico/latest/getting-started/kubernetes/) for the current operator manifest, and replace `CALICO_VERSION` with the latest stable release:

```bash
kubectl create -f https://raw.githubusercontent.com/projectcalico/calico/CALICO_VERSION/manifests/tigera-operator.yaml
kubectl apply -f - <<EOF
apiVersion: operator.tigera.io/v1
kind: Installation
metadata:
  name: default
spec:
  calicoNetwork:
    ipPools:
    - blockSize: 26
      cidr: 10.244.0.0/16
      encapsulation: VXLAN
      natOutgoing: Enabled
EOF
```

Pod IPs change to the new Calico overlay range. Clusters that ran legacy Azure CNI also lose their VNET pod addresses, so update all firewall rules that referenced pod IPs accordingly.

### Service type LoadBalancer

Remove AKS-specific annotations. The cloud controller assigns a public endpoint to a Kubernetes Service with `type: LoadBalancer`.

### Azure monitor DaemonSets

Remove `omsagent` and `ama-logs` DaemonSets before deploying Prometheus; running both causes metrics duplication. Also remove the `azure-policy` pod.

### Container images: ACR to permanent registry

If using Azure Container Registry with AKS-integrated authentication, mirror images to Docker Hub or Harbor. ACR auth tied to AKS managed identity won't work from Quake AI.

### Azure DevOps pipelines

If deploying via Azure DevOps, create a new Kubernetes service connection pointing to the Quake AI cluster:

1. Create a self-hosted agent pool on VMs with access to the Quake AI K8s API
2. Create a kubeconfig-based K8s Service Connection
3. Update pipeline YAML to reference the new connection:

```yaml
- task: KubernetesManifest@1
  inputs:
    action: deploy
    connectionType: kubernetesServiceConnection
    kubernetesServiceConnection: 'quake-ai-k8s'
    manifests: deployment.yaml
```

## 6. Apply and validate

**Migrate data with Velero:**

```bash
# On AKS: install Velero with Azure backend
velero install \
  --provider azure \
  --plugins velero/velero-plugin-for-microsoft-azure:v1.9.0 \
  --bucket aks-migration \
  --use-node-agent \
  --default-volumes-to-fs-backup

velero backup create aks-full --include-namespaces production

# On Quake AI: restore with StorageClass remapping
kubectl apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/plugin-config: ""
    velero.io/change-storage-class: RestoreItemAction
data:
  managed-csi: cinder-flash
  managed-csi-premium: cinder-flash
  azurefile-csi: cinder-flash
EOF

velero restore create --from-backup aks-full --restore-volumes=true
```

**Deploy and verify:**

```bash
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces
kubectl get ingress --all-namespaces
```

For Azure Files NFS data, use `rsync` to the NFS server on Quake AI. For critical databases, prefer application-level dump/restore.

## 7. Observability setup

Replace Azure Monitor Container Insights, Azure Log Analytics, and Application Insights:

```bash
helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set promtail.enabled=true \
  --set loki.persistence.storageClassName=cinder-flash
```

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the `kube-prometheus-stack` chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin `--version 72.9.1`, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the `--version` flag to install the current chart.

For APM traces (replacing Application Insights), deploy the OpenTelemetry Collector with Grafana Tempo.


Remove the `omsagent` / `ama-logs` DaemonSets **before** deploying Prometheus to avoid metrics duplication and resource contention.


## 8. Auto-scaling alternatives

| AKS feature | Quake AI equivalent |
|---|---|
| VMSS Cluster Autoscaler | Cluster Autoscaler with OpenStack provider |
| AKS Node Auto-Provisioning (NAP; GA 2025) | Not available; use Cluster Autoscaler with OpenStack provider or manual Nova scaling |
| KEDA (managed) | KEDA: fully portable via Helm |
| VPA (AKS add-on) | VPA: fully portable |
| HPA | Fully portable |

Azure DevOps can continue deploying to the Quake AI cluster using a self-hosted agent and kubeconfig-based service connection.

## 9. Validation checklist

- [ ] All pods in Running or Completed state
- [ ] PVCs bound to Cinder volumes with correct data
- [ ] Ingress routes working through the ingress controller's public `LoadBalancer` Service
- [ ] DNS resolves to the new Service address
- [ ] cert-manager certificates issued
- [ ] No `azure.workload.identity` annotations remaining
- [ ] No AzureIdentity/AzureIdentityBinding CRDs remaining
- [ ] No SecretProviderClass resources remaining
- [ ] All Key Vault secrets migrated to K8s Secrets or Vault
- [ ] No `omsagent`/`ama-logs` DaemonSets running
- [ ] No `azure-policy` pods running
- [ ] Prometheus scraping all targets
- [ ] Loki ingesting logs
- [ ] Azure DevOps pipeline deploying to Quake AI (if applicable)
- [ ] StatefulSet data integrity verified
- [ ] Application health checks passing

## 10. Provider-specific gotchas

| Gotcha | Impact | Mitigation |
|---|---|---|
| Microsoft Entra Workload Identity OIDC tokens are AKS-specific | Cannot exchange tokens outside Azure | Replace with service principal credentials or migrate services |
| AAD Pod Identity NMI DaemonSet intercepts IMDS | Fails on non-AKS clusters (no IMDS endpoint) | Remove before restore; intercepted pods will fail on startup |
| Azure Key Vault CSI requires Key Vault + Azure AD connectivity | Pod startup fails without Azure network access | Export all secrets to K8s Secrets before migration |
| Azure CNI uses VNET secondary IPs | Pod IPs tied to Azure VNET | New CNI on Quake AI means new pod CIDR; update all firewall rules |
| Azure Files (SMB) has no K8s-native equivalent | SMB volumes are not portable | Migrate to NFS or restructure the application |
| AGIC: one Application Gateway handles multiple clusters | Complex routing config tied to Azure AG | Replicate with nginx-ingress IngressClass |
| `AzureIdentity` CRDs restored by Velero are non-functional | Velero restores CRDs but they can't function | Exclude from backup or delete after restore |
| Azure DevOps managed service connection auth | AKS auth won't work for Quake AI | Create kubeconfig-based service connection |
| Azure Monitor DaemonSets conflict with Prometheus | Running both causes metrics duplication | Remove omsagent before deploying kube-prometheus-stack |
| AKS with Istio add-on (ASM) | Uses AKS-managed Istio control plane | Uninstall; install upstream Istio; re-apply VirtualService/DestinationRule |
| AKS Node Auto-Provisioning (NAP; `AKSNodeClass` and `NodePool` CRDs) | Azure-specific; fails on Quake AI | Remove all NAP CRDs before restore |

## See also

- [Kubernetes migration overview](/docs/kubernetes/migration): all provider guides and portability matrix
- [Migrate from Azure](/resources/migration/from-azure): cross-service Azure migration hub
- [Network migration from Azure VNet](/docs/network/migration/migrate-from-azure-vnet): networking-specific migration
- [Compute migration from Azure VMs](/docs/compute/migration/migrate-from-azure-vms): VM-level migration
