Kubernetes cluster troubleshooting
Kubernetes cluster troubleshooting
This runbook covers two common failure modes for the Kubernetes service (Magnum) on Quake AI: clusters that never leave CREATE_IN_PROGRESS, and kubectl sessions that cannot reach the API server even when the cluster reports success.
Quick triage#
| Symptom | Most likely cause | Jump to |
|---|---|---|
Cluster status stays CREATE_IN_PROGRESS for 30+ minutes with no visible progress | Heat stack failure, WaitCondition timeout, quota or capacity limits, or network/DNS blocking node bootstrap | Cluster stuck in create in progress |
kubectl get nodes returns "Unable to connect to the server" or times out; cluster is CREATE_COMPLETE | Stale kubeconfig, missing or wrong API floating IP, security group blocking TCP 6443, or certificate mismatch | Kubectl connection failure |
Cluster stuck in create in progress#
Symptoms#
The Kubernetes service reports your cluster in CREATE_IN_PROGRESS for longer than 30 minutes. The console shows no state change, and workloads are not available.
Diagnosis#
Quake AI provisions Magnum clusters through the Automation service (Heat
1. Resolve the cluster's Heat stack ID:
openstack coe cluster show YOUR_CLUSTER -f value -c stack_id2. Inspect the stack:
openstack stack show YOUR_STACK_IDNote stack_status. If the stack is CREATE_FAILED, Magnum may still show CREATE_IN_PROGRESS until it reconciles.
3. List failed stack resources:
openstack stack resource list YOUR_STACK_ID --filter status=FAILEDIdentify the resource name and resource_status_reason. A failed WaitCondition means a node-level service (for example etcd, kubelet, or the Kubernetes API server) did not signal readiness in time.
4. If a WaitCondition or node resource failed, check the affected instance:
Open the instance console in the Compute UI or fetch the console log. Look for cloud-init errors, metadata timeouts, image pull failures, or DNS errors.
5. If the stack is CREATE_COMPLETE but Magnum still shows CREATE_IN_PROGRESS:
You are seeing a sync issue between Magnum and Heat. Wait a few minutes and run openstack coe cluster show YOUR_CLUSTER again. If the status does not update, treat the cluster as stuck.
6. Rule out quota and capacity:
Clusters consume multiple instances, volumes, ports, and often floating IPs. Verify project quotas and availability zone capacity. Scheduling failures sometimes surface only in Heat or Nova events, not in the cluster summary.
7. Rule out network and DNS for nodes:
Nodes must reach the metadata service and any registries referenced by the cluster template. Confirm your project network provides working DNS nameservers on the subnet. Without DNS, image pulls and bootstrap scripts fail and WaitConditions time out.
Cluster API (CAPI) drivers are not available on Quake AI until the Epoxy release cycle; production clusters use the Heat driver. Upstream deprecation of the Heat driver does not change current behavior on Antelope-based platforms.
Resolution#
Work from the stack status you observed:
-
Stack
CREATE_FAILED: Fix the root cause shown inresource_status_reason(quota, flavor, image, network, or bootstrap). Update quotas or templates, then delete the failed stack or cluster and recreate. -
WaitCondition or node bootstrap failure: After you fix networking, DNS, or image access, delete the stuck cluster and create a new one with the same template.
-
Stack
CREATE_COMPLETE, Magnum stillCREATE_IN_PROGRESS: Wait for reconciliation. If it persists, delete the cluster and recreate, or open a support ticket with the evidence in Collecting evidence for support tickets.
Delete a stuck cluster:
openstack coe cluster delete YOUR_CLUSTERConfirm deletion completes, then create a new cluster when the underlying issue is resolved.
Verification#
openstack coe cluster show YOUR_CLUSTER -c status -c health_status
openstack stack show YOUR_STACK_ID -c stack_statusBoth Magnum and Heat should show terminal success states before you run kubectl against the cluster.
Prevention#
- Confirm sufficient quota for every resource the template provisions (instances, volumes, ports, floating IPs).
- Use a cluster template you have already validated in the same project and network.
- Set reliable DNS nameservers on subnets used by cluster nodes.
- Keep images and registry endpoints reachable from the node network before you scale production workloads.
When to escalate#
Escalate if the Heat stack succeeds, quotas are adequate, DNS and routing look correct, and the cluster remains CREATE_IN_PROGRESS, or if stack failures reference platform resources you cannot modify. Include cluster name, stack ID, failed resource names, and console log excerpts in the ticket.
Kubectl connection failure#
Symptoms (Kubectl connection failure)#
kubectl get nodes (or any API call) returns Unable to connect to the server, connection timeouts, or TLS errors. openstack coe cluster show reports CREATE_COMPLETE for the same cluster.
Diagnosis (Kubectl connection failure)#
1. Pull a fresh kubeconfig from Magnum:
openstack coe cluster config YOUR_CLUSTER --dir ~/2. Point kubectl at that file:
export KUBECONFIG=~/configStale or hand-edited kubeconfig files are the most common cause of false outages.
3. Read the API server endpoint:
Open ~/config and locate the server: URL under clusters. It must use the public floating IP or hostname that reaches the control plane, not an unreachable private address from your workstation.
4. Confirm the control plane allows TCP 6443 from your source IP:
List security groups attached to the master instances or load balancer fronting the API (per your template). You need an ingress rule for TCP port 6443 from your office IP or bastion, or a defined remote CIDR that includes you.
5. Test raw HTTPS reachability:
curl -k https://API_SERVER_IP:6443/versionA JSON major / minor response means the API is up and reachable. Timeouts mean routing, floating IP, or security group issues. Certificate warnings from curl -k are expected; kubectl uses the CA embedded in the kubeconfig.
6. Confirm master instances are running:
openstack server list | grep YOUR_CLUSTERIf masters are SHUTOFF or in ERROR, the API endpoint will not respond regardless of kubeconfig quality.
Resolution (Kubectl connection failure)#
-
Expired or wrong kubeconfig: Keep using the
openstack coe cluster configoutput; regenerate after any floating IP or endpoint change. -
API not on a reachable IP: Ensure your cluster template assigns a floating IP to the master (or a load balancer with a public VIP). Associate a floating IP if your design expects direct access.
-
Security group blocks 6443: Add an ingress rule for TCP 6443 from your IP or trusted CIDR to the cluster's API security group.
-
Masters down: Start or rebuild instances per your operational playbook; if the cluster is corrupted, delete and recreate the cluster.
Verification (Kubectl connection failure)#
export KUBECONFIG=~/config
kubectl get nodes
kubectl cluster-infoBoth commands should return without connection errors.
Prevention (Kubectl connection failure)#
- Run
openstack coe cluster configafter every create or network change, before you rely onkubectl. - Standardize templates that include a public endpoint (floating IP or load balancer) for the Kubernetes API.
- Document the security group that protects the API and keep TCP 6443 rules aligned with your access model.
When to escalate (Kubectl connection failure)#
Escalate if TCP 6443 is open, masters are healthy, the kubeconfig server URL matches the published floating IP, and curl -k https://API_SERVER_IP:6443/version still fails, or if kubectl reports certificate errors after a fresh config pull. Attach kubeconfig server URL (redact secrets), security group IDs, and curl verbose output.
Collecting evidence for support tickets#
Gather this information before you open a ticket for either failure mode:
| Item | Command |
|---|---|
| Cluster ID, name, and status | openstack coe cluster show YOUR_CLUSTER |
| Associated Heat stack ID and status | openstack coe cluster show YOUR_CLUSTER -f value -c stack_id then openstack stack show YOUR_STACK_ID |
| Failed stack resources | openstack stack resource list YOUR_STACK_ID --filter status=FAILED |
| Cluster nodes (instances) | Run openstack server list and find rows whose name includes your cluster identifier |
| Fresh kubeconfig server URL (no secrets) | After openstack coe cluster config YOUR_CLUSTER --dir ~/, show the server line from ~/config |
| API reachability test | curl -kv https://API_SERVER_IP:6443/version |
| Security group rules for API/masters | openstack security group rule list YOUR_SECURITY_GROUP |
| Console log from a stuck master or minion | openstack console log show YOUR_INSTANCE_NAME --lines 100 |
| Time of first failure (UTC) | Note manually |
For formatting and policy, follow the Support ticket evidence procedure.
See also#
- API error reference: cross-service HTTP status codes and error handling
- Create a Kubernetes cluster: complete cluster creation
- Advanced Kubernetes troubleshooting: deeper cluster-level debugging
- Instance connectivity troubleshooting: SSH, floating IPs, and security groups
- Support ticket evidence procedure: what to attach when you escalate
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.
Last validated: 08.09.2026
Quick answers
- Why does `openstack coe cluster create` fail with a Keystone trust or unauthorized error when I use an application credential?CLIAPITerraform
- Why does a Kubernetes LoadBalancer service stay `<pending>` for several minutes?CLI
- Why does my GitHub Actions or GitLab CI job fail to run `openstack coe` or `kubectl` on a Magnum cluster?CLI
- Why does my Magnum cluster create fail with "Only volume-backed servers" or "Quota exceeded for compute_units"?CLIAPI