# Kubernetes cluster troubleshooting

Source: https://docs.quake.ai/docs/operate/runbooks/kubernetes-troubleshooting
Markdown: https://docs.quake.ai/docs/operate/runbooks/kubernetes-troubleshooting.md
> Diagnose Kubernetes cluster creation failures, stuck pods, networking issues, and storage problems on Quake AI.

---

# Kubernetes cluster troubleshooting

This runbook covers two common failure modes for the Kubernetes service ([Magnum](https://docs.openstack.org/magnum/)) on Quake AI: clusters that never leave `CREATE_IN_PROGRESS`, and `kubectl` sessions that cannot reach the API server even when the cluster reports success.

<Figure size="md" caption="Where to start: pick the branch that matches your symptom, then jump to that section below.">

```d2
direction: right

start: "Kubernetes\ncluster issue"

q_status: "Cluster status?" {shape: diamond}

leaf_build: "Cluster stuck in\ncreate in progress"
leaf_kubectl: "Kubectl connection\nfailure"

start -> q_status
q_status -> leaf_build: CREATE_IN_PROGRESS
q_status -> leaf_kubectl: CREATE_COMPLETE
```

</Figure>

## Quick triage

| Symptom | Most likely cause | Jump to |
|---|---|---|
| Cluster status stays `CREATE_IN_PROGRESS` for 30+ minutes with no visible progress | Heat stack failure, WaitCondition timeout, quota or capacity limits, or network/DNS blocking node bootstrap | [Cluster stuck in create in progress](#cluster-stuck-in-create-in-progress) |
| `kubectl get nodes` returns "Unable to connect to the server" or times out; cluster is `CREATE_COMPLETE` | Stale kubeconfig, missing or wrong API floating IP, security group blocking TCP 6443, or certificate mismatch | [Kubectl connection failure](#kubectl-connection-failure) |

---

## Cluster stuck in create in progress

### Symptoms

The Kubernetes service reports your cluster in `CREATE_IN_PROGRESS` for longer than 30 minutes. The console shows no state change, and workloads are not available.

### Diagnosis

Quake AI provisions Magnum clusters through the Automation service ([Heat](https://docs.openstack.org/heat/)) on Antelope-class releases. Cluster status in Magnum can lag behind the underlying stack, so you always validate the Heat stack before you assume the control plane is still building.

**1. Resolve the cluster's Heat stack ID:**

```bash
openstack coe cluster show YOUR_CLUSTER -f value -c stack_id
```

**2. Inspect the stack:**

```bash
openstack stack show YOUR_STACK_ID
```

Note `stack_status`. If the stack is `CREATE_FAILED`, Magnum may still show `CREATE_IN_PROGRESS` until it reconciles.

**3. List failed stack resources:**

```bash
openstack stack resource list YOUR_STACK_ID --filter status=FAILED
```

Identify the resource name and `resource_status_reason`. A failed WaitCondition means a node-level service (for example etcd, kubelet, or the Kubernetes API server) did not signal readiness in time.

**4. If a WaitCondition or node resource failed, check the affected instance:**

Open the instance console in the Compute UI or fetch the console log. Look for cloud-init errors, metadata timeouts, image pull failures, or DNS errors.

**5. If the stack is `CREATE_COMPLETE` but Magnum still shows `CREATE_IN_PROGRESS`:**

You are seeing a sync issue between Magnum and Heat. Wait a few minutes and run `openstack coe cluster show YOUR_CLUSTER` again. If the status does not update, treat the cluster as stuck.

**6. Rule out quota and capacity:**

Clusters consume multiple instances, volumes, ports, and often floating IPs. Verify project quotas and availability zone capacity. Scheduling failures sometimes surface only in Heat or Nova events, not in the cluster summary.

**7. Rule out network and DNS for nodes:**

Nodes must reach the metadata service and any registries referenced by the cluster template. Confirm your project network provides working DNS nameservers on the subnet. Without DNS, image pulls and bootstrap scripts fail and WaitConditions time out.

Cluster API (CAPI) drivers are not available on Quake AI until the Epoxy release cycle; production clusters use the Heat driver. Upstream deprecation of the Heat driver does not change current behavior on Antelope-based platforms.

### Resolution

Work from the stack status you observed:

- **Stack `CREATE_FAILED`:** Fix the root cause shown in `resource_status_reason` (quota, flavor, image, network, or bootstrap). Update quotas or templates, then delete the failed stack or cluster and recreate.

- **WaitCondition or node bootstrap failure:** After you fix networking, DNS, or image access, delete the stuck cluster and create a new one with the same template.

- **Stack `CREATE_COMPLETE`, Magnum still `CREATE_IN_PROGRESS`:** Wait for reconciliation. If it persists, delete the cluster and recreate, or open a support ticket with the evidence in [Collecting evidence for support tickets](#collecting-evidence-for-support-tickets).

**Delete a stuck cluster:**



`openstack coe cluster delete` permanently destroys the cluster, its nodes, and all workloads running on them. Persistent volumes backed by Cinder survive deletion, but any data on ephemeral storage is lost. Ensure you have backed up critical workloads and configuration before proceeding.



```bash
openstack coe cluster delete YOUR_CLUSTER
```

Confirm deletion completes, then create a new cluster when the underlying issue is resolved.

### Verification

```bash
openstack coe cluster show YOUR_CLUSTER -c status -c health_status
openstack stack show YOUR_STACK_ID -c stack_status
```

Both Magnum and Heat should show terminal success states before you run `kubectl` against the cluster.

### Prevention

- Confirm sufficient quota for every resource the template provisions (instances, volumes, ports, floating IPs).
- Use a cluster template you have already validated in the same project and network.
- Set reliable DNS nameservers on subnets used by cluster nodes.
- Keep images and registry endpoints reachable from the node network before you scale production workloads.

### When to escalate

Escalate if the Heat stack succeeds, quotas are adequate, DNS and routing look correct, and the cluster remains `CREATE_IN_PROGRESS`, or if stack failures reference platform resources you cannot modify. Include cluster name, stack ID, failed resource names, and console log excerpts in the ticket.

---

## Kubectl connection failure

### Symptoms (Kubectl connection failure)

`kubectl get nodes` (or any API call) returns `Unable to connect to the server`, connection timeouts, or TLS errors. `openstack coe cluster show` reports `CREATE_COMPLETE` for the same cluster.

### Diagnosis (Kubectl connection failure)

**1. Pull a fresh kubeconfig from Magnum:**

```bash
openstack coe cluster config YOUR_CLUSTER --dir ~/
```

**2. Point `kubectl` at that file:**

```bash
export KUBECONFIG=~/config
```

Stale or hand-edited kubeconfig files are the most common cause of false outages.

**3. Read the API server endpoint:**

Open `~/config` and locate the `server:` URL under `clusters`. It must use the public floating IP or hostname that reaches the control plane, not an unreachable private address from your workstation.

**4. Confirm the control plane allows TCP 6443 from your source IP:**

List security groups attached to the master instances or load balancer fronting the API (per your template). You need an ingress rule for TCP port `6443` from your office IP or bastion, or a defined remote CIDR that includes you.

**5. Test raw HTTPS reachability:**

```bash
curl -k https://API_SERVER_IP:6443/version
```

A JSON `major` / `minor` response means the API is up and reachable. Timeouts mean routing, floating IP, or security group issues. Certificate warnings from `curl -k` are expected; `kubectl` uses the CA embedded in the kubeconfig.

**6. Confirm master instances are running:**

```bash
openstack server list | grep YOUR_CLUSTER
```

If masters are `SHUTOFF` or in `ERROR`, the API endpoint will not respond regardless of kubeconfig quality.

### Resolution (Kubectl connection failure)

- **Expired or wrong kubeconfig:** Keep using the `openstack coe cluster config` output; regenerate after any floating IP or endpoint change.

- **API not on a reachable IP:** Ensure your cluster template assigns a floating IP to the master (or a load balancer with a public VIP). Associate a floating IP if your design expects direct access.

- **Security group blocks 6443:** Add an ingress rule for TCP 6443 from your IP or trusted CIDR to the cluster's API security group.

- **Masters down:** Start or rebuild instances per your operational playbook; if the cluster is corrupted, delete and recreate the cluster.

### Verification (Kubectl connection failure)

```bash
export KUBECONFIG=~/config
kubectl get nodes
kubectl cluster-info
```

Both commands should return without connection errors.

### Prevention (Kubectl connection failure)

- Run `openstack coe cluster config` after every create or network change, before you rely on `kubectl`.
- Standardize templates that include a public endpoint (floating IP or load balancer) for the Kubernetes API.
- Document the security group that protects the API and keep TCP 6443 rules aligned with your access model.

### When to escalate (Kubectl connection failure)

Escalate if TCP 6443 is open, masters are healthy, the kubeconfig server URL matches the published floating IP, and `curl -k https://API_SERVER_IP:6443/version` still fails, or if `kubectl` reports certificate errors after a fresh config pull. Attach kubeconfig `server` URL (redact secrets), security group IDs, and `curl` verbose output.

---

## Collecting evidence for support tickets

Gather this information before you open a ticket for either failure mode:

| Item | Command |
|---|---|
| Cluster ID, name, and status | `openstack coe cluster show YOUR_CLUSTER` |
| Associated Heat stack ID and status | `openstack coe cluster show YOUR_CLUSTER -f value -c stack_id` then `openstack stack show YOUR_STACK_ID` |
| Failed stack resources | `openstack stack resource list YOUR_STACK_ID --filter status=FAILED` |
| Cluster nodes (instances) | Run `openstack server list` and find rows whose name includes your cluster identifier |
| Fresh kubeconfig server URL (no secrets) | After `openstack coe cluster config YOUR_CLUSTER --dir ~/`, show the `server` line from `~/config` |
| API reachability test | `curl -kv https://API_SERVER_IP:6443/version` |
| Security group rules for API/masters | `openstack security group rule list YOUR_SECURITY_GROUP` |
| Console log from a stuck master or minion | `openstack console log show YOUR_INSTANCE_NAME --lines 100` |
| Time of first failure (UTC) | Note manually |

For formatting and policy, follow the [Support ticket evidence procedure](/docs/operate/troubleshooting/support-ticket-evidence).

## See also

- [API error reference](/reference/errors): cross-service HTTP status codes and error handling
- [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster): complete cluster creation
- [Advanced Kubernetes troubleshooting](/docs/operate/troubleshooting/advanced-kubernetes): deeper cluster-level debugging
- [Instance connectivity troubleshooting](/docs/operate/runbooks/instance-connectivity): SSH, floating IPs, and security groups
- [Support ticket evidence procedure](/docs/operate/troubleshooting/support-ticket-evidence): what to attach when you escalate
