# Migrate from AWS EKS to Kubernetes on Quake AI

Source: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-eks
Markdown: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-eks.md

---

# Migrate from AWS EKS to Kubernetes on Quake AI

## Service mapping

<MigrationTable provider="aws" service="kubernetes" />


## 1. Overview

Amazon EKS gives you a managed Kubernetes control plane with deep AWS integration: VPC-native pod networking (VPC CNI), IAM Roles for Service Accounts (IRSA), EBS/EFS-backed persistent storage, ALB/NLB ingress through the AWS Load Balancer Controller, and Karpenter for sub-minute node autoscaling.

EKS also offers Auto Mode (generally available since December 2024), which manages both the control plane and the worker nodes through Karpenter, running EC2 managed instances. Auto Mode clusters add managed CRDs (`NodePool` and `EC2NodeClass`). EKS Pod Identity (generally available since November 2023) is the newer pod-to-IAM mechanism: it uses `PodIdentityAssociation` resources instead of the IRSA `eks.amazonaws.com/role-arn` annotation. If your source cluster uses Auto Mode or Pod Identity, add the matching cleanup items to the checklist in Section 3.

On Quake AI, you provision Kubernetes two ways. You can run **self-managed RKE2 or k3s on Nova instances**, installing and operating the control plane yourself and wiring in the OpenStack cloud-provider stack (CCM + Cinder CSI), or you can provision a cluster through [Magnum](/docs/kubernetes), the Container Infrastructure Management service, which builds the cluster from a curated template (`openstack coe cluster create`). Both are supported; see the [Kubernetes migration overview](/docs/kubernetes/migration) for the trade-offs.

For an EKS migration, self-managed often makes it easier to match your source cluster's Kubernetes version and CNI exactly: the Magnum template catalog provides a fixed set of versions that can differ from the versions EKS supports (1.30 through 1.35 as of June 2026). Run `openstack coe cluster template list` to see the templates available in your region before you choose Magnum.

**Migration complexity: Medium-High.** The primary drivers are IRSA removal, EBS-to-Cinder data migration, ALB-to-nginx-ingress conversion, VPC CNI replacement with Calico/Cilium, Fargate pod migration to regular node pools, and Karpenter replacement. Expect **4-8 weeks** for a production migration depending on IRSA and EFS usage.

## 2. Workload portability matrix

| Resource type | Portability | Notes |
|---|---|---|
| Deployment | Portable | No changes needed |
| StatefulSet (manifest) | Portable | You migrate data separately through Velero |
| DaemonSet | Portable | You may need toleration or selector adjustments |
| ConfigMap | Portable | You update any AWS endpoint values |
| Secret (values) | Portable | Re-encode if sourced from AWS Secrets Manager |
| Service (ClusterIP, NodePort) | Portable | No changes |
| Service (LoadBalancer) | Needs adaptation | Remove AWS load-balancer annotations; the cloud controller assigns the public endpoint |
| Ingress | Needs adaptation | Replace ALB annotations with nginx-ingress |
| PersistentVolumeClaim | Needs adaptation | Change `storageClassName` from `gp2`/`gp3` to `cinder-flash` |
| PersistentVolume (data) | Not portable | You migrate data through Velero filesystem backup |
| NetworkPolicy | Portable | You need Calico or Cilium on Quake AI |
| HPA / VPA / PDB | Portable | No changes |
| CronJob / Job | Portable | No changes |
| ServiceAccount | Needs adaptation | Remove `eks.amazonaws.com/role-arn` annotations |
| RBAC resources | Portable | No changes |
| AWS VPC CNI (`aws-node`) | Provider-specific | Replace with Calico or Cilium |
| ALB Controller | Provider-specific | Replace with ingress-nginx or Traefik and a `LoadBalancer` Service |
| Karpenter NodePool/EC2NodeClass | Provider-specific | Remove; use Cluster Autoscaler or manual scaling |
| Fargate profiles | Provider-specific | Deploy to regular node pools |
| EFS CSI (`efs.csi.aws.com`) | Provider-specific | No direct equivalent; use NFS server or Rook-Ceph for RWX |

## 3. Pre-migration: export and audit

Audit your existing EKS cluster to identify all provider-specific resources before you make changes.

Confirm your EKS kubeconfig is current and that `kubectl` can reach the source cluster:

```bash
aws eks update-kubeconfig --region REGION --name CLUSTER_NAME
kubectl get nodes
```

A modern EKS kubeconfig uses `aws eks get-token` as the exec credential plugin. You no longer need the older `aws-iam-authenticator` binary with AWS CLI 1.16.156 or later.

```bash
kubectl get all --all-namespaces -o yaml > cluster-export.yaml

kubectl get pvc --all-namespaces -o yaml > pvcs.yaml

kubectl get sa --all-namespaces -o yaml | grep -l 'eks.amazonaws.com' > irsa-accounts.txt

kubectl get ingress --all-namespaces -o yaml > ingresses.yaml

helm list --all-namespaces > helm-releases.txt

kubectl get pods --all-namespaces -o json | \
  jq '.items[] | select(.metadata.annotations["eks.amazonaws.com/compute-type"]=="fargate") | .metadata.name'
```

**Inventory checklist:**

- [ ] All EBS-backed PVCs and their StorageClasses (`gp2`, `gp3`, custom)
- [ ] All IRSA-annotated ServiceAccounts (`eks.amazonaws.com/role-arn`)
- [ ] All ALB/NLB Ingress resources and their annotations
- [ ] All Fargate-scheduled pods (`eks.amazonaws.com/compute-type: fargate`)
- [ ] All ECR image references
- [ ] EKS Pod Identity associations (`PodIdentityAssociation` resources, in addition to IRSA-annotated ServiceAccounts)
- [ ] EKS Auto Mode NodePools and `EC2NodeClass` CRDs (if applicable)
- [ ] Helm releases and their versions
- [ ] EFS PVCs (RWX): these require special handling

## 4. Provision Kubernetes on Quake AI

### 4.1 Option A: self-managed RKE2 or k3s (recommended)

Build the cluster yourself on Nova instances. Use OpenTofu with the `openstack` provider to create the infrastructure, then install Kubernetes. This path matches any Kubernetes version, CNI, or CRI your workloads require, including the versions EKS supports today.

**Infrastructure (OpenTofu):**
- Neutron network + subnet
- Security groups (allow K8s API on port 6443, NodePort range 30000-32767, inter-node communication)
- Nova instances: 3 control plane nodes (for HA) + N worker nodes
- Stable endpoint for the Kubernetes API

**Install RKE2 (one supported self-managed distribution):**

```bash
curl -sfL https://get.rke2.io | INSTALL_RKE2_TYPE=server sh -
systemctl enable --now rke2-server
```

**Install the OpenStack cloud-provider stack on the self-managed cluster:**
1. Deploy `openstack-cloud-controller-manager` with a `cloud.conf` using [application credentials](/docs/identity/how-to/create-application-credential)
2. Deploy the Cinder CSI driver with the `cinder-flash` StorageClass
3. Install Calico (VXLAN mode) or Cilium as the CNI

Verify the cluster:

```bash
kubectl get nodes
kubectl get csinodes
kubectl get storageclass
```

### 4.2 Option B: Magnum


**Magnum cluster creation requires a password-scoped session.** `openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Trust delegation is not available to Keystone application credentials, so the call fails before any Magnum work begins. Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` (password auth) before running the command. See [the Kubernetes FAQ](/docs/kubernetes/faq) for the full troubleshooting flow.


Provision a cluster from a Magnum template. The platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API. Run `openstack coe cluster template list` first to confirm which templates and Kubernetes versions your region offers.


**Magnum cluster creation requires a password-scoped session.** `openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call the infrastructure APIs they need during lifecycle operations. Trust delegation is not available to Keystone application credentials, so the call fails before any Magnum work begins. Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` before running the command. The Cloud Console wizard works from any logged-in user session. See [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster) for details.


```bash
openstack coe cluster template list
openstack coe cluster create production-k8s \
  --cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
  --master-count 3 \
  --node-count 3 \
  --keypair MY_KEYPAIR
```

For the full walkthrough including Console and quota guidance, see [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster). To roll your own template (custom flavors, alternate Kubernetes version), see [Create a cluster template](/docs/kubernetes/how-to/create-cluster-template).

Once the cluster reports `CREATE_COMPLETE`, fetch the kubeconfig:

```bash
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclass
```

The Cinder CSI driver and the OpenStack cloud-provider stack are pre-installed by the template; no manual CCM, CSI, or CNI install is required for the Magnum path.

## 5. Adapt provider-specific resources

### CSI driver: EBS to Cinder

Replace all references to the EBS CSI driver and its StorageClasses:

```yaml
# Before (EKS)
storageClassName: gp2

# After (Quake AI)
storageClassName: cinder-flash
```

The `cinder-flash` StorageClass uses the Cinder CSI provisioner with `Flash_Premium` volume type:

```yaml
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: cinder-flash
  annotations:
    storageclass.kubernetes.io/is-default-class: "true"
provisioner: cinder.csi.openstack.org
parameters:
  type: Flash_Premium
allowVolumeExpansion: true
volumeBindingMode: WaitForFirstConsumer
```

**EFS (RWX) workloads:** Amazon EFS gives you ReadWriteMany access with no direct Cinder equivalent. Your options on Quake AI:
- NFS server provisioner (dedicated Nova VM running NFS)
- Rook-Ceph with CephFS (production-grade RWX, operationally complex)
- Restructure the application to use RWO volumes with application-level coordination

### Data migration with Velero

Use [Velero](https://velero.io/) to back up your EKS workloads and restore them onto the Quake AI cluster, including filesystem-level backup of PersistentVolume data. Install the Velero provider plugin that matches your Velero version instead of pinning an old release:

```bash
velero plugin add velero/velero-plugin-for-aws:CURRENT_VERSION
```

Use the plugin version compatible with your Velero installation; see the [Velero compatibility matrix](https://github.com/vmware-tanzu/velero).

To remap EBS StorageClasses to `cinder-flash` during the restore, create a `change-storage-class` ConfigMap before you run `velero restore create`:

```yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/change-storage-class: RestoreItemAction
data:
  gp2: cinder-flash
  gp3: cinder-flash
```

Create this ConfigMap before running `velero restore create`. The `change-storage-class` RestoreItemAction is built into Velero core (Velero 1.10 and later). You still need provider plugins for the backup storage location, but not for StorageClass remapping.

### Ingress controller: ALB to nginx-ingress

Replace AWS Load Balancer Controller annotations with nginx-ingress annotations:

```yaml
# Before (EKS ALB)
annotations:
  kubernetes.io/ingress.class: alb
  alb.ingress.kubernetes.io/scheme: internet-facing
  alb.ingress.kubernetes.io/target-type: ip

# After (Quake AI nginx-ingress)
annotations:
  kubernetes.io/ingress.class: nginx
  nginx.ingress.kubernetes.io/ssl-redirect: "true"
```

TLS termination moves from ACM (AWS Certificate Manager) to cert-manager with Let's Encrypt on Quake AI. You deploy cert-manager and create a `ClusterIssuer` before you migrate Ingress resources.

### Pod identity: IRSA to application credentials

IRSA injects AWS credentials into pods by annotating ServiceAccounts with `eks.amazonaws.com/role-arn`. This mechanism does not work outside AWS.

EKS Pod Identity (generally available since November 2023) is an alternative to IRSA that uses `PodIdentityAssociation` resources instead of ServiceAccount annotations. If your source cluster uses Pod Identity, list the associations during the audit and remove the EKS Pod Identity add-on before migration:

```bash
aws eks list-pod-identity-associations --cluster-name CLUSTER_NAME
```

**For applications accessing OpenStack APIs:** Create [application credentials](/docs/identity/how-to/create-application-credential) and store them in Kubernetes Secrets:

```bash
openstack application credential create MYAPP_CREDENTIAL --role member
kubectl create secret generic openstack-app-creds \
  --from-literal=OS_APPLICATION_CREDENTIAL_ID=APPLICATION_CREDENTIAL_ID \
  --from-literal=OS_APPLICATION_CREDENTIAL_SECRET=APPLICATION_CREDENTIAL_SECRET
```

Replace `APPLICATION_CREDENTIAL_ID` and `APPLICATION_CREDENTIAL_SECRET` with the values from the `openstack application credential create` output.

**For applications that still need AWS access (hybrid period):** Replace IRSA with static IAM credentials in Kubernetes Secrets. This is less secure than IRSA but works during the transition.

You remove the IRSA annotation from every affected ServiceAccount:

```bash
kubectl annotate sa MY_SERVICE_ACCOUNT eks.amazonaws.com/role-arn- -n MY_NAMESPACE
```

### Service type LoadBalancer

Kubernetes Services with `type: LoadBalancer` receive a public endpoint through the cluster's cloud controller. Remove AWS-specific annotations and keep the standard Service fields. The assigned address appears in `status.loadBalancer.ingress`.

### Fargate pods

Remove the `eks.amazonaws.com/compute-type: fargate` annotation and any Fargate profile references. Deploy these workloads to regular Nova-backed worker nodes. Note the isolation change: Fargate enforces hard resource isolation per pod (dedicated vCPU and memory), while Quake AI node pools enforce limits through `cgroups` and let pods share node resources. Apply `LimitRange` resources to namespaces to enforce minimum resource requests and prevent unbounded consumption.

### Container images: ECR to permanent registry

ECR authentication tokens expire every 12 hours and cannot be used from outside AWS long-term. Mirror all images to a permanent registry (Docker Hub, self-hosted Harbor, or Quay.io) before cutover:

```bash
docker pull AWS_ACCOUNT_ID.dkr.ecr.AWS_REGION.amazonaws.com/IMAGE_NAME:IMAGE_TAG
docker tag AWS_ACCOUNT_ID.dkr.ecr.AWS_REGION.amazonaws.com/IMAGE_NAME:IMAGE_TAG docker.io/DOCKER_ORG/IMAGE_NAME:IMAGE_TAG
docker push docker.io/DOCKER_ORG/IMAGE_NAME:IMAGE_TAG
```

Replace `AWS_ACCOUNT_ID`, `AWS_REGION`, `IMAGE_NAME`, `IMAGE_TAG`, and `DOCKER_ORG` with your registry values. You update all Deployment and StatefulSet manifests with the new image references.

## 6. Apply and validate

You deploy adapted manifests to the Quake AI cluster:

```bash
kubectl apply -f adapted-manifests/ --namespace production
kubectl get pods -n production -w
```

**Verify each layer:**

```bash
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces
kubectl get ingress --all-namespaces
```

Test application endpoints through the ingress controller's public `LoadBalancer` Service. Confirm PVCs bind to Cinder volumes.

## 7. Observability setup

You replace AWS managed observability (CloudWatch Container Insights, X-Ray, CloudWatch Logs) with a self-managed stack:

```bash
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

helm repo add grafana https://grafana.github.io/helm-charts
helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set promtail.enabled=true \
  --set loki.persistence.enabled=true \
  --set loki.persistence.storageClassName=cinder-flash
```

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the `kube-prometheus-stack` chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin `--version 72.9.1`, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the `--version` flag to install the current chart.

The Helm releases install Prometheus (metrics), Grafana (dashboards), Alertmanager (alerting), and Loki with Promtail (log aggregation). For distributed tracing, you add Grafana Tempo to replace X-Ray.

## 8. Auto-scaling alternatives

| EKS feature | Quake AI equivalent |
|---|---|
| Karpenter | Not available. Karpenter (`karpenter.sh/v1` stable API, including `EC2NodeClass`) is EKS-specific; use Cluster Autoscaler with the OpenStack provider. EKS Auto Mode uses Karpenter internally. Remove all Karpenter CRDs before Velero restore. |
| EKS Auto Mode (GA December 2024) | Not available. Auto Mode nodes are fully managed through Karpenter. Remove the Auto Mode `NodePool` and `EC2NodeClass` CRDs before migration. |
| Cluster Autoscaler (ASG) | Cluster Autoscaler with OpenStack Nova API |
| Fargate (serverless pods) | Regular Nova-backed node pools; use resource quotas |
| Managed node groups | OpenTofu-managed Nova instance groups + `kubeadm join` |

HPA and VPA stay fully portable and behave the same on Quake AI.

For the Cluster Autoscaler on OpenStack, you set the `--cloud-provider=openstack` flag and define node groups by Nova server groups or Cluster API (CAPO).

## 9. Validation checklist

- [ ] All pods in Running or Completed state
- [ ] PVCs bound to Cinder volumes with correct data
- [ ] Ingress routes return expected responses through the public `LoadBalancer` Service
- [ ] DNS resolves to the new Service address
- [ ] Prometheus scraping all targets (check Targets page in Grafana)
- [ ] Loki ingesting logs from all namespaces
- [ ] cert-manager certificates issued and valid
- [ ] Application health checks passing
- [ ] StatefulSet data integrity verified (database queries, object counts)
- [ ] No IRSA-related errors in pod logs
- [ ] No EBS/ALB/ECR references remaining in manifests

## 10. Provider-specific gotchas

| Gotcha | Impact | Mitigation |
|---|---|---|
| Fargate pods use `eks.amazonaws.com/compute-type` annotation | Pods won't schedule without Fargate on non-EKS clusters | Remove annotation; deploy to regular node pool |
| ECR auth tokens expire every 12 hours | Quake AI cluster can't pull ECR images after expiry | Mirror images to a permanent registry before cutover |
| IRSA OIDC endpoint is EKS-cluster-specific | Token exchange only works within AWS | Replace with explicit credentials or migrate to OpenStack services |
| VPC CNI pod IPs are VPC-routable | Cross-cluster pod routing assumptions break | New CIDR range on Quake AI; update firewall rules accordingly |
| EBS volumes are AZ-pinned | Can't move volume data between AZs | Migrate data via Velero; Cinder volumes are region-level |
| ALB annotations are extensive (WAF, ACM, routing rules) | Complex annotation translation required | Audit each ALB; TLS termination moves to cert-manager + nginx |
| Karpenter `EC2NodeClass` CRDs | Validation errors on non-EKS clusters | Remove all Karpenter CRDs and NodePools before Velero restore |
| AWS CNI Security Group Policies | SGPs are EKS/AWS-only | Replace with K8s NetworkPolicy on Calico/Cilium |
| EFS (RWX) has no direct Cinder equivalent | Applications requiring shared storage need restructuring | NFS server provisioner, Rook-Ceph, or application-level sharding |

## See also

- [Kubernetes migration overview](/docs/kubernetes/migration): all provider guides and portability matrix
- [Migrate from AWS](/resources/migration/from-aws): cross-service AWS migration hub
- [Network migration from AWS VPC](/docs/network/migration/migrate-from-aws-vpc): networking-specific migration
- [Compute migration from EC2](/docs/compute/migration/migrate-from-ec2): VM-level migration
