Skip to content

Kubernetes cluster troubleshooting

Runbook · Updated Sep 2026

Kubernetes cluster troubleshooting

This runbook covers two common failure modes for the Kubernetes service (Magnum) on Quake AI: clusters that never leave CREATE_IN_PROGRESS, and kubectl sessions that cannot reach the API server even when the cluster reports success.

Kubernetescluster issueCluster status?Cluster stuck increate in progressKubectl connectionfailure CREATE_IN_PROGRESSCREATE_COMPLETE
Click to zoom
Where to start: pick the branch that matches your symptom, then jump to that section below.

Quick triage#

SymptomMost likely causeJump to
Cluster status stays CREATE_IN_PROGRESS for 30+ minutes with no visible progressHeat stack failure, WaitCondition timeout, quota or capacity limits, or network/DNS blocking node bootstrapCluster stuck in create in progress
kubectl get nodes returns "Unable to connect to the server" or times out; cluster is CREATE_COMPLETEStale kubeconfig, missing or wrong API floating IP, security group blocking TCP 6443, or certificate mismatchKubectl connection failure

Cluster stuck in create in progress#

Symptoms#

The Kubernetes service reports your cluster in CREATE_IN_PROGRESS for longer than 30 minutes. The console shows no state change, and workloads are not available.

Diagnosis#

Quake AI provisions Magnum clusters through the Automation service (Heat) on Antelope-class releases. Cluster status in Magnum can lag behind the underlying stack, so you always validate the Heat stack before you assume the control plane is still building.

1. Resolve the cluster's Heat stack ID:

bash
openstack coe cluster show YOUR_CLUSTER -f value -c stack_id

2. Inspect the stack:

bash
openstack stack show YOUR_STACK_ID

Note stack_status. If the stack is CREATE_FAILED, Magnum may still show CREATE_IN_PROGRESS until it reconciles.

3. List failed stack resources:

bash
openstack stack resource list YOUR_STACK_ID --filter status=FAILED

Identify the resource name and resource_status_reason. A failed WaitCondition means a node-level service (for example etcd, kubelet, or the Kubernetes API server) did not signal readiness in time.

4. If a WaitCondition or node resource failed, check the affected instance:

Open the instance console in the Compute UI or fetch the console log. Look for cloud-init errors, metadata timeouts, image pull failures, or DNS errors.

5. If the stack is CREATE_COMPLETE but Magnum still shows CREATE_IN_PROGRESS:

You are seeing a sync issue between Magnum and Heat. Wait a few minutes and run openstack coe cluster show YOUR_CLUSTER again. If the status does not update, treat the cluster as stuck.

6. Rule out quota and capacity:

Clusters consume multiple instances, volumes, ports, and often floating IPs. Verify project quotas and availability zone capacity. Scheduling failures sometimes surface only in Heat or Nova events, not in the cluster summary.

7. Rule out network and DNS for nodes:

Nodes must reach the metadata service and any registries referenced by the cluster template. Confirm your project network provides working DNS nameservers on the subnet. Without DNS, image pulls and bootstrap scripts fail and WaitConditions time out.

Cluster API (CAPI) drivers are not available on Quake AI until the Epoxy release cycle; production clusters use the Heat driver. Upstream deprecation of the Heat driver does not change current behavior on Antelope-based platforms.

Resolution#

Work from the stack status you observed:

  • Stack CREATE_FAILED: Fix the root cause shown in resource_status_reason (quota, flavor, image, network, or bootstrap). Update quotas or templates, then delete the failed stack or cluster and recreate.

  • WaitCondition or node bootstrap failure: After you fix networking, DNS, or image access, delete the stuck cluster and create a new one with the same template.

  • Stack CREATE_COMPLETE, Magnum still CREATE_IN_PROGRESS: Wait for reconciliation. If it persists, delete the cluster and recreate, or open a support ticket with the evidence in Collecting evidence for support tickets.

Delete a stuck cluster:

bash
openstack coe cluster delete YOUR_CLUSTER

Confirm deletion completes, then create a new cluster when the underlying issue is resolved.

Verification#

bash
openstack coe cluster show YOUR_CLUSTER -c status -c health_status
openstack stack show YOUR_STACK_ID -c stack_status

Both Magnum and Heat should show terminal success states before you run kubectl against the cluster.

Prevention#

  • Confirm sufficient quota for every resource the template provisions (instances, volumes, ports, floating IPs).
  • Use a cluster template you have already validated in the same project and network.
  • Set reliable DNS nameservers on subnets used by cluster nodes.
  • Keep images and registry endpoints reachable from the node network before you scale production workloads.

When to escalate#

Escalate if the Heat stack succeeds, quotas are adequate, DNS and routing look correct, and the cluster remains CREATE_IN_PROGRESS, or if stack failures reference platform resources you cannot modify. Include cluster name, stack ID, failed resource names, and console log excerpts in the ticket.


Kubectl connection failure#

Symptoms (Kubectl connection failure)#

kubectl get nodes (or any API call) returns Unable to connect to the server, connection timeouts, or TLS errors. openstack coe cluster show reports CREATE_COMPLETE for the same cluster.

Diagnosis (Kubectl connection failure)#

1. Pull a fresh kubeconfig from Magnum:

bash
openstack coe cluster config YOUR_CLUSTER --dir ~/

2. Point kubectl at that file:

bash
export KUBECONFIG=~/config

Stale or hand-edited kubeconfig files are the most common cause of false outages.

3. Read the API server endpoint:

Open ~/config and locate the server: URL under clusters. It must use the public floating IP or hostname that reaches the control plane, not an unreachable private address from your workstation.

4. Confirm the control plane allows TCP 6443 from your source IP:

List security groups attached to the master instances or load balancer fronting the API (per your template). You need an ingress rule for TCP port 6443 from your office IP or bastion, or a defined remote CIDR that includes you.

5. Test raw HTTPS reachability:

bash
curl -k https://API_SERVER_IP:6443/version

A JSON major / minor response means the API is up and reachable. Timeouts mean routing, floating IP, or security group issues. Certificate warnings from curl -k are expected; kubectl uses the CA embedded in the kubeconfig.

6. Confirm master instances are running:

bash
openstack server list | grep YOUR_CLUSTER

If masters are SHUTOFF or in ERROR, the API endpoint will not respond regardless of kubeconfig quality.

Resolution (Kubectl connection failure)#

  • Expired or wrong kubeconfig: Keep using the openstack coe cluster config output; regenerate after any floating IP or endpoint change.

  • API not on a reachable IP: Ensure your cluster template assigns a floating IP to the master (or a load balancer with a public VIP). Associate a floating IP if your design expects direct access.

  • Security group blocks 6443: Add an ingress rule for TCP 6443 from your IP or trusted CIDR to the cluster's API security group.

  • Masters down: Start or rebuild instances per your operational playbook; if the cluster is corrupted, delete and recreate the cluster.

Verification (Kubectl connection failure)#

bash
export KUBECONFIG=~/config
kubectl get nodes
kubectl cluster-info

Both commands should return without connection errors.

Prevention (Kubectl connection failure)#

  • Run openstack coe cluster config after every create or network change, before you rely on kubectl.
  • Standardize templates that include a public endpoint (floating IP or load balancer) for the Kubernetes API.
  • Document the security group that protects the API and keep TCP 6443 rules aligned with your access model.

When to escalate (Kubectl connection failure)#

Escalate if TCP 6443 is open, masters are healthy, the kubeconfig server URL matches the published floating IP, and curl -k https://API_SERVER_IP:6443/version still fails, or if kubectl reports certificate errors after a fresh config pull. Attach kubeconfig server URL (redact secrets), security group IDs, and curl verbose output.


Collecting evidence for support tickets#

Gather this information before you open a ticket for either failure mode:

ItemCommand
Cluster ID, name, and statusopenstack coe cluster show YOUR_CLUSTER
Associated Heat stack ID and statusopenstack coe cluster show YOUR_CLUSTER -f value -c stack_id then openstack stack show YOUR_STACK_ID
Failed stack resourcesopenstack stack resource list YOUR_STACK_ID --filter status=FAILED
Cluster nodes (instances)Run openstack server list and find rows whose name includes your cluster identifier
Fresh kubeconfig server URL (no secrets)After openstack coe cluster config YOUR_CLUSTER --dir ~/, show the server line from ~/config
API reachability testcurl -kv https://API_SERVER_IP:6443/version
Security group rules for API/mastersopenstack security group rule list YOUR_SECURITY_GROUP
Console log from a stuck master or minionopenstack console log show YOUR_INSTANCE_NAME --lines 100
Time of first failure (UTC)Note manually

For formatting and policy, follow the Support ticket evidence procedure.

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

For the full policy, see Usage Guidelines.

Last validated: 08.09.2026

Quick answers

Was this page helpful?