# Migrate from DigitalOcean DOKS to Kubernetes on Quake AI

Source: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-doks
Markdown: https://docs.quake.ai/docs/kubernetes/migration/migrate-from-doks.md

---

# Migrate from DigitalOcean DOKS to Kubernetes on Quake AI

## Service mapping

<MigrationTable provider="do" service="kubernetes" />


## 1. Overview

DigitalOcean Kubernetes (DOKS) provides a managed control plane with minimal vendor lock-in. DOKS uses vanilla Kubernetes with DO-specific add-ons: the `dobs.csi.digitalocean.com` CSI driver for block storage, the DigitalOcean cloud controller manager for load balancers, and optional DO Container Registry for images. DOKS offers an optional high-availability (HA) control plane; on Quake AI you set control-plane redundancy yourself through the Magnum `--master-count` or the control-plane node count of a self-managed cluster.

On Quake AI, you provision Kubernetes through [Magnum](/docs/kubernetes), the Container Infrastructure Management service. Magnum gives you a control plane and worker nodes from a curated cluster template (`openstack coe cluster create`). The region ships default cluster templates, each pre-wired with Calico CNI, the OpenStack cloud-provider stack (CCM + Cinder CSI), and a load-balanced API endpoint. Run `openstack coe cluster template list` to see the Kubernetes versions currently available. DOKS tracks the latest three upstream minor versions, so compare that list against your source cluster and plan for a version jump if the Magnum catalog trails your DOKS version. DOKS users will recognize this as a familiar managed-control-plane shape.

Quake AI supports two provisioning paths: a Magnum cluster from a curated template, or **self-managed RKE2 or k3s on Nova instances**. Self-managed is the better fit when you need a custom CNI, custom kubelet flags, custom CRI, or a Kubernetes version outside the template catalog. Both paths are covered below.

**Migration complexity: Low.** DOKS has no IAM-for-pods pattern, no proprietary CNI, and limited vendor-specific annotations. The primary tasks are swapping the CSI driver/CCM and mirroring container images. Expect **1-2 weeks** for a typical production migration.

## 2. Workload portability matrix

| Resource type | Portability | Notes |
|---|---|---|
| Deployment | Portable | No changes needed |
| StatefulSet (manifest) | Portable | Data migrated separately via Velero |
| DaemonSet | Portable | No changes needed |
| ConfigMap | Portable | Update any DO endpoint values |
| Secret (values) | Portable | Remove `DIGITALOCEAN_ACCESS_TOKEN` references |
| Service (ClusterIP, NodePort) | Portable | No changes |
| Service (LoadBalancer) | Needs adaptation | Remove `do-loadbalancer-*` annotations |
| Ingress | Portable | nginx-ingress is commonly used on DOKS already |
| PersistentVolumeClaim | Needs adaptation | Change `storageClassName` from `do-block-storage` to `cinder-flash` |
| PersistentVolume (data) | Not portable | Migrate data via Velero filesystem backup |
| NetworkPolicy | Portable | DOKS uses Cilium internally; Calico/Cilium on Quake AI |
| HPA / VPA / PDB | Portable | No changes |
| CronJob / Job | Portable | No changes |
| RBAC resources | Portable | No changes |
| DO CSI driver (`dobs.csi.digitalocean.com`) | Provider-specific | Replace with Cinder CSI |
| DO cloud controller (`digitalocean`) | Provider-specific | Replace with OpenStack CCM |
| DO Container Registry images | Provider-specific | Mirror to Docker Hub, Harbor, or Quay.io |

## 3. Pre-migration: export and audit

```bash
kubectl get all --all-namespaces -o yaml > cluster-export.yaml

kubectl get pvc,pv --all-namespaces -o yaml > storage.yaml

kubectl get svc --all-namespaces -o yaml | grep 'do-loadbalancer' > do-lb-services.txt

helm list --all-namespaces > helm-releases.txt
```

**Inventory checklist:**

- [ ] All PVCs using `do-block-storage` StorageClass
- [ ] All Services with `service.beta.kubernetes.io/do-loadbalancer-*` annotations
- [ ] Images hosted on `registry.digitalocean.com`
- [ ] Helm releases and versions
- [ ] external-dns configuration (if using DO DNS)

## 4. Provision Kubernetes on Quake AI

### 4.1 Magnum

Provision a cluster from a Magnum template. The platform creates the control plane VMs, worker VMs, security groups, networking, and a stable endpoint for the Kubernetes API. The Kubernetes version comes from the template catalog; if you need a version, CNI, or CRI it does not cover, use the self-managed path in Section 4.2.


**Magnum cluster creation requires a password-scoped session.** `openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Trust delegation is not available to Keystone application credentials, so the call fails before any Magnum work begins. Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` (password auth) before running the command. The Cloud Console wizard works from any logged-in user session because it uses the user's password-scoped token. See [the Kubernetes FAQ](/docs/kubernetes/faq) for the full troubleshooting flow.


```bash
openstack coe cluster template list
openstack coe cluster create production-k8s \
  --cluster-template Standard-v2.0-k8s-calico-fc38_v1.24.16 \
  --master-count 3 \
  --node-count 3 \
  --keypair MY_KEYPAIR
```

For the full walkthrough including Console and quota guidance, see [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster). To roll your own template (custom flavors, alternate Kubernetes version), see [Create a cluster template](/docs/kubernetes/how-to/create-cluster-template).

Once the cluster reports `CREATE_COMPLETE`, fetch the kubeconfig:

```bash
openstack coe cluster config production-k8s --dir ~/.kube
kubectl get nodes
kubectl get storageclass
```

The Cinder CSI driver and the OpenStack cloud-provider stack are pre-installed by the template; no manual CCM, CSI, or CNI install is required for the Magnum path.

### 4.2 Self-managed RKE2 or k3s (recommended)

If the Magnum template catalog does not match your version, CNI, or CRI constraints, build the cluster yourself on Nova instances. Use OpenTofu with the `openstack` provider to create the infrastructure:

- Neutron network + subnet
- Security groups (K8s API 6443, NodePort 30000-32767, inter-node)
- Nova instances: 3 control plane + N workers
- Stable endpoint for the Kubernetes API

Install RKE2 or k3s (both are commonly used on DOKS as well), then deploy on the self-managed cluster:
1. `openstack-cloud-controller-manager` with [application credentials](/docs/identity/how-to/create-application-credential)
2. Cinder CSI driver with the `cinder-flash` StorageClass
3. Calico (VXLAN) or Cilium as the CNI

## 5. Adapt provider-specific resources

### CSI driver: DO block storage to Cinder

Replace the StorageClass reference in all PVCs:

```yaml
# Before (DOKS)
storageClassName: do-block-storage

# After (Quake AI)
storageClassName: cinder-flash
```

DOKS volumes are standard RWO block storage, a direct 1:1 replacement with Cinder. No topology constraints or AZ-pinning complexity.

### Ingress controller

If your DOKS cluster already uses nginx-ingress (the most common pattern), **no ingress changes are needed** beyond DNS updates. nginx-ingress is fully portable.

If using DO-managed load balancers directly, remove DO-specific annotations:

```yaml
# Remove these DO-specific annotations
service.beta.kubernetes.io/do-loadbalancer-name
service.beta.kubernetes.io/do-loadbalancer-protocol
service.beta.kubernetes.io/do-loadbalancer-size-slug
```

The cloud controller assigns a public endpoint to a Kubernetes Service with `type: LoadBalancer` without requiring provider-specific annotations.

### Pod identity

DOKS has no IAM-for-pods mechanism. Pods on DOKS access DigitalOcean APIs using long-lived tokens in K8s Secrets.

- Remove any `DIGITALOCEAN_ACCESS_TOKEN` secret references
- For applications that called DO APIs (e.g., external-dns with DO DNS provider), reconfigure to use the target DNS provider:

```yaml
# Before: external-dns with DO DNS
args:
  - --provider=digitalocean

# After: external-dns with Cloudflare (or other)
args:
  - --provider=cloudflare
```

### Service type LoadBalancer

The standard Kubernetes `LoadBalancer` Service fields remain portable. Remove any `service.beta.kubernetes.io/do-loadbalancer-*` annotations.

### Container images: DO container registry to permanent registry

DO Container Registry is DO-specific. Mirror all images before cutover:

```bash
docker pull registry.digitalocean.com/REGISTRY/IMAGE:TAG
docker tag registry.digitalocean.com/REGISTRY/IMAGE:TAG docker.io/ORG/IMAGE:TAG
docker push docker.io/ORG/IMAGE:TAG
```

Update all Deployment and StatefulSet manifests with new image references. This is the most manual step in a DOKS migration.

## 6. Apply and validate

**Migrate data with Velero:**

```bash
# On DOKS: install Velero with S3/Spaces backend
velero install \
  --provider aws \
  --plugins velero/velero-plugin-for-aws:v1.9.0 \
  --bucket doks-migration \
  --use-node-agent \
  --default-volumes-to-fs-backup \
  --backup-location-config \
    region=nyc3,s3ForcePathStyle=true,s3Url=https://nyc3.digitaloceanspaces.com

velero backup create doks-full --include-namespaces production

# On Quake AI: restore with StorageClass remapping
kubectl apply -f - <<EOF
apiVersion: v1
kind: ConfigMap
metadata:
  name: change-storage-class-config
  namespace: velero
  labels:
    velero.io/plugin-config: ""
    velero.io/change-storage-class: RestoreItemAction
data:
  do-block-storage: cinder-flash
EOF

velero restore create --from-backup doks-full --restore-volumes=true
```

**Deploy and verify:**

```bash
kubectl get pods --all-namespaces -o wide
kubectl get pvc --all-namespaces
kubectl get svc --all-namespaces
```

For stateful workloads (databases), prefer application-level dump/restore (pg_dump, mysqldump) over Velero filesystem backup for guaranteed consistency.

## 7. Observability setup

Most DOKS users already run self-managed Prometheus + Grafana. If so, these migrate as-is; only PVC StorageClass references need updating.

If starting fresh, install the self-managed observability stack:

```bash
helm upgrade --install kube-prometheus-stack \
  prometheus-community/kube-prometheus-stack \
  --version 72.9.1 \
  --namespace monitoring --create-namespace

helm upgrade --install loki grafana/loki-stack \
  --namespace monitoring \
  --set promtail.enabled=true \
  --set loki.persistence.storageClassName=cinder-flash
```

Match the chart version to your cluster's Kubernetes version. From chart version 73.0.0 the `kube-prometheus-stack` chart requires Kubernetes 1.25 or later. Quake AI Magnum cluster templates run Kubernetes 1.24.16, so pin `--version 72.9.1`, the last chart release that supports 1.24. On a cluster running Kubernetes 1.25 or later, omit the `--version` flag to install the current chart.

DigitalOcean's managed metrics (basic node/cluster metrics in the DO console) have no equivalent on Quake AI. All monitoring is self-managed.

## 8. Auto-scaling alternatives

| DOKS feature | Quake AI equivalent |
|---|---|
| Managed node autoscaling (DO API) | Cluster Autoscaler with OpenStack provider, or manual Nova scaling |
| HPA | Fully portable |
| VPA | Fully portable |

DOKS node autoscaling supports min/max bounds per node pool, and GPU node pools can scale to zero. That maps to similar operational complexity on the OpenStack Cluster Autoscaler or manual Nova scaling.

## 9. Validation checklist

- [ ] All pods in Running or Completed state
- [ ] PVCs bound to Cinder volumes with correct data
- [ ] Ingress routes working through the ingress controller's public `LoadBalancer` Service
- [ ] DNS resolves to new floating IPs
- [ ] Prometheus scraping all targets
- [ ] No `do-loadbalancer-*` annotations remaining
- [ ] No `registry.digitalocean.com` image references remaining
- [ ] No `DIGITALOCEAN_ACCESS_TOKEN` secrets remaining
- [ ] StatefulSet data integrity verified
- [ ] Application health checks passing

## 10. Provider-specific gotchas

| Gotcha | Impact | Mitigation |
|---|---|---|
| DO Container Registry tokens expire | Quake AI can't pull DOCR images | Mirror all images before migration |
| DO-specific load-balancer annotations do not apply | The Service may not receive the expected configuration | Review and remove `do-loadbalancer-*` annotations |
| DOKS auto-applies system-level manifests | DO operator annotations/labels may appear | Strip DO-injected metadata before restoring on Quake AI |
| DO volumes capped at 7 per node | May need to rethink node sizing | Cinder has no per-node volume limit (subject to project quota) |
| DOKS internal Cilium is transparent | NetworkPolicy works on DOKS via Cilium but behavior may differ | Test NetworkPolicy on Quake AI's Calico/Cilium to confirm identical behavior |
| VPC-native DOKS (1.31+) pods are VPC-routable | Service integrations or firewall rules that reach pod IPs directly stop working | On Quake AI's Calico VXLAN, pod IPs are overlay-only and not VPC-routable. Update any firewall rules or service integrations that depend on direct pod IP reachability from outside the cluster |

## See also

- [Kubernetes migration overview](/docs/kubernetes/migration): all provider guides and portability matrix
- [Migrate from DigitalOcean](/resources/migration/from-digitalocean): cross-service DO migration hub
- [Network migration from DO VPC](/docs/network/migration/migrate-from-do-vpc): networking-specific migration
- [Object storage migration from Spaces](/docs/object/migration/migrate-from-spaces): storage migration
