Kubernetes FAQ
Kubernetes FAQ
Frequently asked questions about Kubernetes clusters on Quake AI. For step-by-step instructions see the Kubernetes how-to guides. For deeper background see the concept pages. For migration from another provider see the migration guides.
Getting started#
What is the Kubernetes service on Quake AI?
The Kubernetes service automates cluster provisioning on top of existing Compute, Network, and Storage infrastructure. You define the cluster shape through a reusable cluster template, and the platform provisions master VMs, worker VMs, a dedicated cluster network, stable endpoints for the Kubernetes API and etcd, and the router connecting the cluster to the external network. Provisioning is backed by OpenStack Magnum and orchestrated through the Automation service (Heat).
Once the cluster is running, you manage workloads with standard tools: kubectl, Helm, and the OpenStack CLI for cluster-level operations such as scaling or deletion. The platform handles the infrastructure setup; you operate the cluster after provisioning, including control plane availability and upgrades.
Source: Kubernetes service overview, Kubernetes on Quake AI
Is this a managed Kubernetes control plane?
Quake AI Kubernetes provisions the control plane inside your project. The master nodes run as Compute instances, so your team operates control plane availability, etcd backups, and certificate rotation. The platform automates the provisioning of those master VMs and the networking between them.
This model is for teams that want Kubernetes infrastructure provisioning automated without setting up kubeadm, CNI, control-plane endpoints, and storage drivers by hand.
Source: Kubernetes service overview: What this service is (and isn't), Kubernetes on Quake AI: How Quake AI Kubernetes compares
What infrastructure does a cluster provision inside my project?
When you create a cluster, the Kubernetes service provisions the following resources in your project:
| Resource | Role |
|---|---|
| Master VMs | Run the Kubernetes API server, scheduler, controller manager, and etcd. Provisioned as Compute instances. |
| Worker VMs | Run the kubelet and container runtime. Your application pods are scheduled here. |
| Network | A dedicated network and subnet for intra-cluster communication. |
| Control-plane endpoints | A stable Kubernetes API endpoint when master load balancing is enabled. Clusters with multiple master nodes also use a stable etcd endpoint. |
| Router | Connects the cluster network to the external network. |
| Persistent volumes | Backed by Block Storage volumes through the Cinder CSI driver. |
Underlying VMs use Fedora CoreOS images with the container runtime and Kubernetes components pre-installed. The Heat stack that drives provisioning is visible under Automation > Stacks in the console.
Source: Kubernetes on Quake AI: What the Kubernetes service provisions
What are the minimum resource requirements to create a cluster?
Kubernetes clusters consume Compute, Network, and Storage quotas. Minimum requirements by cluster size:
| Cluster size | Master nodes | Worker nodes | Min vCPU | Min RAM | Min block storage | Floating IPs |
|---|---|---|---|---|---|---|
| Minimal (dev/test) | 1 | 1 | 4 shared | 4 GB | 180 GB | 3 |
| Small production | 3 | 2 | 34 dedicated | 80 GB | 500 GB | 4 |
| Medium production | 3 | 5 | 46 dedicated | 128 GB | 1 TB | 7 |
Check your quota before creating a cluster with openstack quota show or navigate to Home > Overview in the console.
Source: Kubernetes on Quake AI: Resource requirements and quotas
Core concepts#
What is a cluster template, and why do I need one?
A cluster template is a reusable blueprint that captures every parameter the Kubernetes service needs to provision a cluster:
- Kubernetes version and container orchestration engine (COE)
- Compute flavors for master and worker nodes
- Network driver (Flannel or Calico) and external network
- Volume driver for persistent storage (Cinder)
- TLS settings, Docker storage driver, and registry configuration
- Labels that control feature flags (autoscaler, dashboard, network policies)
You create a template once and reuse it for multiple clusters. Platform-provided templates (for example, Standard-V2.0-k8s-calico-fc38_v1.24.16) are available by default. Separating the template from the cluster lets you standardize configurations across teams and environments, for instance, a production template with three master nodes and Calico, and a development template with a single master and Flannel.
Source: Kubernetes on Quake AI: Cluster templates, How to create a cluster template
Which network driver should I choose: Calico or Flannel?
Choose based on your need for Kubernetes NetworkPolicy enforcement:
- Calico: recommended for production. Supports Kubernetes
NetworkPolicyfor pod-to-pod traffic control. Configured in VXLAN mode on Quake AI migrations. - Flannel: simpler overlay network with lower operational overhead. Does not enforce
NetworkPolicy. Suitable for development clusters or workloads that do not require pod-level traffic isolation.
The network driver is set in the cluster template and cannot be changed after the cluster is created.
Source: How to create a cluster template: CLI parameters, Kubernetes migration overview: Recommended stack
What does the cluster lifecycle look like?
| Phase | Status | What happens |
|---|---|---|
| Create | CREATE_IN_PROGRESS → CREATE_COMPLETE | The service provisions VMs, networking, and control-plane endpoints from the cluster template. Creation typically takes 5 to 15 minutes. |
| Run | ACTIVE / health status | The cluster is active. Interact with it via kubectl using a kubeconfig retrieved from the CLI. |
| Scale | UPDATE_IN_PROGRESS → UPDATE_COMPLETE | Add or remove worker nodes. The platform provisions or removes VMs accordingly. |
| Update | UPDATE_IN_PROGRESS → UPDATE_COMPLETE | Modify cluster properties such as node count or labels. |
| Delete | DELETE_IN_PROGRESS → DELETE_COMPLETE | The platform tears down the associated VMs, networks, control-plane endpoints, and Heat stack. |
Cluster health is monitored continuously. The console displays both the operational status and a health status based on node and control plane checks.
Source: Kubernetes on Quake AI: Cluster lifecycle, How to manage a Kubernetes cluster
How do I achieve high availability for the Kubernetes control plane?
Create the cluster with 3 master nodes. Quake AI places master VMs on separate physical hosts using Nova anti-affinity scheduling (the same server-group mechanism available for Compute instances). The cluster uses stable, load-balanced endpoints for the Kubernetes API and etcd across the three masters.
For information on how anti-affinity server groups work at the Compute level, see Server groups.
Source: How to create a Kubernetes cluster: Master Node Count, How to manage a Kubernetes cluster: Scale worker nodes, Kubernetes on Quake AI: Resource requirements
What Kubernetes objects and tools are fully portable when migrating?
The following resources migrate without modification: Deployments, StatefulSets, DaemonSets, ConfigMaps, Secrets (values), Namespaces, ServiceAccounts (minus cloud IAM annotations), all RBAC resources, NetworkPolicy, HPA, VPA, PodDisruptionBudget, CronJobs, and Jobs.
Helm charts using standard Kubernetes resources deploy cleanly with updated values. The following operators and Helm charts are also fully portable: cert-manager, external-dns, kube-prometheus-stack, Loki, Grafana, Velero, Istio, Linkerd, ArgoCD, Flux, ingress-nginx, Traefik, Metrics Server, and HashiCorp Vault.
Items that need adaptation: Ingress class and annotations, Service type: LoadBalancer annotations (remove provider-specific ones), and PersistentVolumeClaim StorageClass references (change to cinder-flash).
Source: Kubernetes migration overview: What's portable and what's not
Operations and lifecycle#
How do I create a cluster?
Prerequisites: a cluster template, an SSH key pair, and sufficient project quota.
Console: Services > Kubernetes > Clusters > Create Cluster. Select a template, set master count (use 3 for production), set worker count, configure networking, and optionally enable Auto Healing and Auto Scaling. Creation takes 5 to 15 minutes.
CLI:
openstack coe cluster create MY_CLUSTER_NAME \
--cluster-template MY_TEMPLATE_NAME \
--keypair MY_KEYPAIR \
--master-count 3 \
--node-count 2 \
--timeout 60Monitor status with openstack coe cluster show MY_CLUSTER_NAME -f value -c status. Wait for CREATE_COMPLETE.
Source: How to create a Kubernetes cluster
Why does my Magnum cluster create fail with "Only volume-backed servers" or "Quota exceeded for compute_units"?
Magnum cluster creation on Quake AI requires two adjustments the upstream defaults do not provide: a boot_volume_size label on the cluster template, and flavor choices that fit the project's compute_units quota.
Boot volume size. Every Quake AI flavor has disk: 0. Magnum's upstream default produces image-backed (zero-disk) servers, which the platform rejects with Forbidden: Only volume-backed servers are allowed for flavors with zero disk. Add boot_volume_size=40 (the platform-template default) to the cluster template's --labels argument:
openstack coe cluster template create MY_TEMPLATE_NAME \
--coe kubernetes \
--image FedoraCoreOS-38 \
--keypair MY_KEYPAIR \
--flavor c2a.large \
--master-flavor c2a.xlarge \
--external-network PublicStatic \
--network-driver calico \
--volume-driver cinder \
--master-lb-enabled \
--labels boot_volume_size=40,master_lb_floating_ip_enabled=trueCompute units quota. The platform's compute_units quota caps how many Compute units a project can provision at once. Default-tier projects start at 8000. Each flavor carries a per-instance reservation in its rumble:compute_units property:
| Flavor | vCPUs | RAM | compute_units |
|---|---|---|---|
c2a.large | 2 | 4 GB | 2000 |
c2a.xlarge | 4 | 8 GB | 4000 |
c2a.2xlarge | 8 | 16 GB | 8000 |
A cluster with a single c2a.xlarge master and a single c2a.large worker consumes 6000 compute_units, fitting the default quota. A c2a.2xlarge master alone consumes 8000 compute_units and leaves no headroom for any other resource in the project. Verify your quota with openstack quota show -f value -c compute_units before sizing, and request an increase through the Compute Service Request flow before going larger.
The platform-provided cluster templates (for example, Standard-V2.0-k8s-calico-fc38_v1.24.16) already carry the four required labels and a working flavor sizing. Cloning a platform template is the fastest path to a working cluster.
Source: How to create a cluster template, How to create a Kubernetes cluster, Flavors
Why does `openstack coe cluster create` fail with a Keystone trust or unauthorized error when I use an application credential?
Magnum cluster creation requires a password-scoped session. openstack coe cluster create triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Keystone application credentials cannot delegate trust, so the call fails before any Magnum work begins. The cluster transitions to CREATE_FAILED with this error:
Resource CREATE failed: Forbidden: resources.kube_cluster_deploy:
Failed to create trustee or trust for ClusterAuthenticate the CLI session with OS_USERNAME and OS_PASSWORD (password auth) before running openstack coe cluster create:
export OS_AUTH_URL=https://keystone.us-east-1.rumble.cloud/v3
export OS_PROJECT_ID=YOUR_PROJECT_ID
export OS_USERNAME=YOUR_USERNAME
export OS_PASSWORD=YOUR_PASSWORD
export OS_USER_DOMAIN_NAME=Default
export OS_PROJECT_DOMAIN_NAME=Default
export OS_IDENTITY_API_VERSION=3
openstack coe cluster create ...The Console wizard at Services > Kubernetes > Clusters > Create Cluster works from any logged-in user session because it uses the user's password-scoped token. The same constraint applies to Terraform and Heat templates that drive Magnum: configure the OpenStack provider with password auth before applying a template that creates a cluster.
Application credentials remain the right choice for other workloads on Quake AI, including Compute, Network, Volume, Object Storage, and DNS. The Magnum trust-delegation requirement is the documented exception.
Source: How to create a Kubernetes cluster, How to create a cluster template, How to create application credentials
How do I scale worker nodes?
Scaling only applies to worker nodes. Master node count is fixed at creation time.
Console: Services > Kubernetes > Clusters > Resize Cluster > enter the new node count > OK. The cluster status moves to UPDATE_IN_PROGRESS while VMs are provisioned or removed.
CLI:
openstack coe cluster resize MY_CLUSTER_NAME NODE_COUNTMonitor with openstack coe cluster show MY_CLUSTER_NAME -f value -c status until status returns to UPDATE_COMPLETE.
Source: How to manage a Kubernetes cluster: Scale worker nodes
How do I upgrade Kubernetes versions?
The Kubernetes service does not support in-place version upgrades. The recommended approach is:
- Create a new cluster template that specifies the target Kubernetes version.
- Provision a new cluster from that template.
- Migrate workloads to the new cluster using Velero or by redeploying manifests.
- Delete the old cluster once traffic is shifted and data is confirmed intact.
This approach keeps the old cluster available for rollback during the migration window.
Source: Kubernetes on Quake AI: Cluster upgrades
How do I access my cluster with kubectl?
Use the CLI to retrieve the kubeconfig file after the cluster reaches CREATE_COMPLETE:
openstack coe cluster config MY_CLUSTER_NAMEThis outputs an export KUBECONFIG=... command. Run it, then verify connectivity:
kubectl get nodesAll master and worker nodes should show Ready status. The console does not provide direct kubectl access; use the CLI method.
Source: How to create a Kubernetes cluster: Verify the result, How to manage a Kubernetes cluster: Access the cluster with kubectl
How do I delete a cluster?
Deleting a cluster removes its VMs, networks, control-plane endpoints, routers, and underlying Heat stack.
Console: Services > Kubernetes > Clusters > Delete > confirm.
CLI:
openstack coe cluster delete MY_CLUSTER_NAMEIf deletion fails with DELETE_FAILED, check openstack coe cluster show MY_CLUSTER_NAME -f value -c status_reason. Common causes include orphaned resources or dependency conflicts.
Source: How to manage a Kubernetes cluster: Delete a cluster
Troubleshooting#
Why is my cluster stuck in CREATE_IN_PROGRESS?
A cluster that stays in CREATE_IN_PROGRESS for longer than 30 minutes typically indicates one of four root causes:
- Heat stack failure: inspect the underlying stack with
openstack coe cluster show YOUR_CLUSTER -f value -c stack_id, thenopenstack stack show YOUR_STACK_ID. If the stack isCREATE_FAILED, readresource_status_reason. - WaitCondition timeout: a node-level service (etcd, kubelet, or the Kubernetes API server) did not signal readiness in time. Check the instance console log for cloud-init errors, metadata timeouts, or image pull failures.
- Quota or capacity limits: clusters consume multiple instances, volumes, ports, and often floating IPs. Verify project quotas with
openstack quota showand check availability zone capacity. - Network or DNS blocking bootstrap: nodes must reach the metadata service and any image registries referenced by the cluster template. Confirm DNS nameservers are set correctly on the cluster subnet.
If the Heat stack is CREATE_COMPLETE but Magnum still shows CREATE_IN_PROGRESS, wait a few minutes for Magnum to reconcile. If it does not update, treat the cluster as stuck and delete it.
Source: Kubernetes cluster troubleshooting: Cluster stuck in create in progress
Why can't kubectl connect to my cluster after CREATE_COMPLETE?
kubectl get nodes returning "Unable to connect to the server" or a timeout despite the cluster showing CREATE_COMPLETE has four common causes:
- Stale or hand-edited kubeconfig: regenerate with
openstack coe cluster config YOUR_CLUSTER --dir ~/. This is the most common cause. - API server not on a reachable IP: open
~/configand verify theserver:URL uses a public floating IP, not an unreachable private address. - Security group blocking TCP 6443: confirm the cluster's API security group allows ingress on TCP port 6443 from your IP or trusted CIDR.
- Master instances down: run
openstack server list | grep YOUR_CLUSTERto confirm masters areACTIVE, notSHUTOFFor inERROR.
Test raw API reachability with curl -k https://API_SERVER_IP:6443/version. A JSON response confirms the API is up; a timeout points to routing, floating IP, or security group issues.
Source: Kubernetes cluster troubleshooting: Kubectl connection failure
Why does my GitHub Actions or GitLab CI job fail to run `openstack coe` or `kubectl` on a Magnum cluster?
CI images often ship python-openstackclient without the Magnum (coe) plugin, or a slim Python base without kubectl and Helm. Install the full client stack in the job before you call cluster APIs:
pip install python-openstackclient python-magnumclientFor GitLab python:3.12-slim and similar images, also install kubectl and Helm from their upstream releases (or use a runner image that already includes them). Authenticate with a password-scoped OpenStack session when the workflow creates or reconfigures clusters; application credentials cannot delegate the Keystone trust Magnum requires.
Source: Deploy from CI, Deploy a Helm chart
Why does a Kubernetes LoadBalancer service stay `<pending>` for several minutes?
A Service with type: LoadBalancer stays <pending> while the cluster provisions its public endpoint and floating IP. Provisioning usually completes within a few minutes; long delays often trace to floating IP quota limits, a subnet without a router to PublicStatic, or a cluster template with floating_ip_enabled: false.
Run kubectl describe svc SERVICE_NAME -n NAMESPACE and check Events for cloud-controller or networking errors. Confirm the project still has floating IP capacity with openstack quota show --usage. Regenerate kubeconfig with openstack coe cluster config CLUSTER_NAME --dir "$HOME/.kube/CLUSTER_NAME" so Helm and kubectl target the cluster you tested.
Source: Deploy a Helm chart, Kubernetes cluster troubleshooting
Migration#
What is the recommended path for migrating Kubernetes workloads to Quake AI?
The migration guides all share a common pattern:
- Audit your source cluster: export all manifests, identify provider-specific resources (cloud CSI drivers, cloud controller manager, ingress controllers, pod identity mechanisms).
- Provision a new self-managed Kubernetes cluster on Nova instances using OpenTofu with the
openstackprovider, then install RKE2 (recommended for production) or k3s. - Install the OpenStack cloud-provider stack:
openstack-cloud-controller-manager, Cinder CSI driver withcinder-flashStorageClass, and Calico (VXLAN) or Cilium as the CNI. - Adapt provider-specific resources: replace the cloud CSI driver with Cinder CSI, replace the cloud ingress controller with nginx-ingress, replace pod-identity mechanisms with OpenStack application credentials in Kubernetes Secrets.
- Migrate data with Velero using
--default-volumes-to-fs-backupand StorageClass remapping tocinder-flash. - Validate workloads and cut over DNS.
Source: Kubernetes migration overview
What is the recommended migration path from AWS EKS?
Complexity: Medium-High. Estimated timeline: 4 to 8 weeks.
Key tasks unique to EKS:
- IRSA removal: remove
eks.amazonaws.com/role-arnannotations from every ServiceAccount; replace with OpenStack application credentials in Kubernetes Secrets. - EBS to Cinder: change
storageClassNamefromgp2/gp3tocinder-flashin all PVCs; migrate data with Velero filesystem backup. - ALB to nginx-ingress: replace ALB annotations with nginx-ingress annotations; TLS termination moves from ACM to cert-manager + Let's Encrypt.
- VPC CNI replacement: replace
aws-nodeDaemonSet with Calico or Cilium. - Fargate pods: remove
eks.amazonaws.com/compute-type: fargateannotations; deploy to regular Nova-backed node pools. - ECR images: mirror all ECR images to Docker Hub or Harbor before cutover; ECR auth tokens expire every 12 hours.
- Karpenter: remove all Karpenter CRDs and NodePools; replace with Cluster Autoscaler with
--cloud-provider=openstackor manual Nova scaling.
See also the Coming from AWS: Kubernetes section for a service-by-service terminology mapping.
Source: Migrate from AWS EKS, Coming from AWS: Kubernetes: EKS → Magnum
What is the recommended migration path from Azure AKS?
Complexity: High. Estimated timeline: 6 to 12 weeks. This is the most complex of the five migrations.
Key tasks unique to AKS:
- Export Azure Key Vault secrets first, pod startup fails without Azure network access, so copy all secrets to Kubernetes Secrets (or HashiCorp Vault) before migrating pods.
- Entra ID Workload Identity / AAD Pod Identity removal: remove
azure.workload.identity/client-idannotations; removeAzureIdentityandAzureIdentityBindingCRDs and the NMI DaemonSet. - AGIC to nginx-ingress: replace Application Gateway Ingress Controller annotations.
- Azure CNI to Calico/Cilium: Azure CNI uses VNET secondary IPs; the new pod CIDR changes, so update all firewall rules.
- Azure Managed Disk CSI to Cinder: change
storageClassNamefrommanaged-csi/managed-csi-premiumtocinder-flash. - Azure Monitor DaemonSets: remove
omsagent/ama-logsDaemonSets before deploying Prometheus to avoid metrics duplication. - Azure DevOps pipelines: create a kubeconfig-based Kubernetes service connection pointing to the Quake AI cluster.
Source: Migrate from Azure AKS
What is the recommended migration path from GCP GKE?
Complexity: Medium-High. Estimated timeline: 3 to 6 weeks (longer if Autopilot or Anthos Service Mesh is in use).
Key tasks unique to GKE:
- Workload Identity Federation removal: remove
iam.gke.io/gcp-service-accountannotations; replace with service account JSON keys (hybrid period) or OpenStack application credentials. - Config Connector CRDs: remove all
cnrm.cloud.google.comCRDs; provision equivalent OpenStack resources via OpenTofu. - GKE Ingress to nginx-ingress: replace
ingressClassName: gcewithingressClassName: nginx; removeManagedCertificate,BackendConfig, andFrontendConfigCRDs; replace with cert-manager. - GCE PD CSI to Cinder: change
storageClassNamefromstandard/standard-rwo/premium-rwotocinder-flash. - Autopilot constraints: re-enable DaemonSets (CCM, CSI, Calico all require them), allow privileged containers, remove Autopilot-specific node tolerations.
- Anthos Service Mesh: uninstall ASM; install upstream Istio or Linkerd; re-apply service mesh policies using upstream CRDs.
Source: Migrate from GCP GKE
What is the recommended migration path from DigitalOcean DOKS?
Complexity: Low. Estimated timeline: 1 to 2 weeks. DOKS is the simplest migration because it has minimal vendor lock-in.
Key tasks:
- DO CSI to Cinder: change
storageClassNamefromdo-block-storagetocinder-flash. - DO CCM to OpenStack CCM: remove
do-loadbalancer-*annotations from Services; the destination cloud controller handles standardLoadBalancerServices. - Ingress: if already using nginx-ingress (common on DOKS), no ingress changes are required beyond DNS updates.
- No pod identity to migrate: DOKS has no IAM-for-pods mechanism; remove
DIGITALOCEAN_ACCESS_TOKENsecret references. - DO Container Registry images: mirror all images to Docker Hub, Harbor, or Quay.io before cutover; DOCR tokens expire.
Source: Migrate from DigitalOcean DOKS
What is the recommended migration path from Hetzner Managed Kubernetes?
Complexity: Low-Medium. Estimated timeline: 1 to 2 weeks. Hetzner's architecture is structurally identical to the self-managed path on Quake AI, swap the cloud provider components and rewrite the OpenTofu IaC.
Key tasks:
- hcloud CSI to Cinder: change
storageClassNamefromhcloud-volumestocinder-flash. - hcloud CCM to OpenStack CCM: replace the
hcloudsecret inkube-systemwith the OpenStackcloud-configsecret; removeload-balancer.hetzner.cloud/*annotations from Services. - IaC rewrite: update OpenTofu from the
hcloudprovider to theopenstackprovider (resource concepts are similar:hcloud_server→openstack_compute_instance_v2,hcloud_network→openstack_networking_network_v2, etc.). - Cluster Autoscaler: swap
--cloud-provider=hcloudfor--cloud-provider=openstack; reconfigure node groups to use Nova server groups. - If using nginx-ingress (common on Hetzner). It migrates with zero ingress changes beyond DNS updates.
Source: Migrate from Hetzner Managed Kubernetes
Pricing, quotas, availability, and support#
What regions and availability zones is the Kubernetes service available in?
The Kubernetes service runs in all three Quake AI regions: us-east-1, us-east-2, and us-west-1. Each cluster lives in a single region; cross-region cluster federation is not provided. Quake AI uses regions, not availability zones, so master and worker placement diversity within a region is expressed through anti-affinity server groups, not through AZ selection. See Regions for the current region table and console URLs.
Source: Regions and availability, Server groups
What is the SLA for the Kubernetes service?
To raise effective control-plane availability, run masters in an anti-affinity server group across at least three Compute instances and back up etcd on a schedule. Master nodes run in your project as Compute instances, so you own control-plane uptime rather than relying on a managed control-plane SLA from EKS, GKE, or AKS. What is covered by the platform SLA is the underlying Compute, Network, and Block Storage primitives the cluster runs on; those contractual terms are published at rumble.cloud/legal and explained on the Service Level Agreement page.
Source: Service Level Agreement, Kubernetes on Quake AI: How Quake AI Kubernetes compares
How do I get help if my cluster is broken?
For most problems, start with the troubleshooting runbook:
- Kubernetes cluster troubleshooting: covers cluster stuck in
CREATE_IN_PROGRESSandkubectlconnection failures, with diagnostic commands and escalation criteria.
Before opening a support ticket, gather the following evidence:
| Item | Command |
|---|---|
| Cluster ID, name, and status | openstack coe cluster show YOUR_CLUSTER |
| Associated Heat stack ID and status | openstack coe cluster show YOUR_CLUSTER -f value -c stack_id, then openstack stack show YOUR_STACK_ID |
| Failed stack resources | openstack stack resource list YOUR_STACK_ID --filter status=FAILED |
| API reachability test | curl -kv https://API_SERVER_IP:6443/version |
| Console log from a stuck master | openstack console log show YOUR_INSTANCE_NAME --lines 100 |
If the runbook does not resolve the issue, open a support ticket with the evidence above. Get help is the canonical decision tree for picking the right starting point based on the symptom, and Support ticket evidence collection is the full evidence checklist.
Source: Get help, Kubernetes cluster troubleshooting: Collecting evidence for support tickets
See also#
- Kubernetes service overview: use cases, guides index, and security considerations
- Kubernetes on Quake AI: architecture, cluster templates, resource planning, and lifecycle
- Cloud-Native Computing on Quake AI: containers vs. VMs, orchestration, CI/CD patterns
- How to create a cluster template
- How to create a Kubernetes cluster
- How to manage a Kubernetes cluster
- Kubernetes cluster troubleshooting: symptom → cause → fix for cluster creation and
kubectlfailures - Kubernetes migration overview: portability matrix, recommended stack, and Velero data migration
- Migrate from AWS EKS
- Migrate from Azure AKS
- Migrate from GCP GKE
- Migrate from DigitalOcean DOKS
- Migrate from Hetzner Managed Kubernetes
- Coming from AWS: Kubernetes: EKS → Magnum
- Server groups: how anti-affinity placement works at the Compute level
- Kubernetes CLI reference
- Kubernetes console reference