# How to Create a Kubernetes Cluster

Source: https://docs.quake.ai/docs/kubernetes/how-to/create-cluster
Markdown: https://docs.quake.ai/docs/kubernetes/how-to/create-cluster.md

---

# How to create a Kubernetes cluster

Deploy a Kubernetes cluster with master and worker nodes from an existing cluster template. The Kubernetes service provisions the VMs, networking, and stable control-plane endpoints, giving you a working cluster you can manage with `kubectl`.

<PrerequisiteBlock methods={["console", "cli"]}>

- A [cluster template](/docs/kubernetes/how-to/create-cluster-template) that defines the Kubernetes version, default flavor, and network configuration
- An [SSH key pair](/docs/tools/add-ssh-key) uploaded to your account
- Sufficient project quota for the planned cluster size. A minimal 1-master + 1-worker cluster on `c2a.xlarge` + `c2a.large` consumes 6000 of the default-tier 8000 `compute_units` budget. See [Sizing and quota](/docs/kubernetes/how-to/create-cluster-template#sizing-and-quota) on the template page for the flavor-to-`compute_units` table; check the live total with `openstack limits show --absolute` before sizing.

</PrerequisiteBlock>


**Floating IP quota.** Magnum clusters consume one floating IP for the public Kubernetes API endpoint and, at upstream defaults, one per master node and one per worker node. A 1-master and 1-worker cluster with `--master-lb-enabled` and node floating IPs allocated needs 3 floating IPs. To minimize demand, the [`iac/templates/k8s-cluster`](/resources/iac-templates/k8s-cluster) template gives the API endpoint one floating IP and disables node floating IPs with `floating_ip_enabled=false`.



**Zero-disk flavors require `boot_volume_size`.** Every Quake AI flavor family ships with `disk: 0`. Magnum cluster creation fails within about 60 seconds with `Forbidden: Only volume-backed servers are allowed for flavors with zero disk (HTTP 403)` unless the cluster's resolved label set carries `boot_volume_size=N` (40 GB minimum). The platform-provided cluster templates carry `boot_volume_size=40` by default; verify on Step 5 that the label is present before selecting **Confirm**. See [Sizing and quota](/docs/kubernetes/how-to/create-cluster-template#sizing-and-quota) on the template page.



**Single-master clusters cannot grow.** A cluster created with `--master-count 1` cannot add masters later. Pass `--master-count 3` at create time for any cluster you intend to keep beyond development.



**Magnum cluster creation requires a password-scoped session.** `openstack coe cluster create` triggers Keystone trust delegation so cluster nodes can call back to OpenStack on your behalf. Trust delegation is not available to Keystone application credentials, so the call fails before any Magnum work begins. Authenticate the CLI session with `OS_USERNAME` and `OS_PASSWORD` (password auth) before running the command. The Cloud Console wizard works from any logged-in user session because it uses the user's password-scoped token. See [the Kubernetes FAQ](/docs/kubernetes/faq) for the full troubleshooting flow.


## Create the cluster

<MethodTabs>
<Method label="Console">

the Console renders the Create Cluster wizard as five numbered steps: **Cluster Info**, **Node Spec**, **Network Setting**, **Management**, and **Additional Labels**. The bottom action bar reads **Cancel | Previous: `<step>` | Next: `<step>`** on steps 1 through 4, and **Cancel | Previous: Management | Confirm** on step 5.


**Kubernetes-only, version set by the template.** Quake AI Magnum runs Kubernetes only. Neither the Create Cluster wizard nor the Cluster Template wizard exposes a COE (Container Orchestration Engine) selector, so the **COE** column on the template table in Step 1 reads `kubernetes`. The template also fixes the Kubernetes version: platform-provided templates use the `FedoraCoreos-38` base image and set the version through the `kube_tag` label. Image selection does not change the Kubernetes version. The bundled `Standard-V2.0-k8s-calico-fc38_*` templates cover the currently supported versions.


1. Go to **Kubernetes** > **Clusters** and select **Create Cluster**.

2. **Step 1: Cluster Info.**
   - **Cluster Name** (required): a name for the cluster (for example, `production-k8s`).
   - **Cluster Template** (required): pick a template from the multi-column table. Columns are **ID/Name**, **COE**, **Network Driver**, and **Keypair**. Selecting a row places a chip below the table reading **Selected: `<template-name>`** with a clear control. Platform-provided templates include the Kubernetes version in the name (for example, `Standard-V2.0-k8s-calico-fc38_v1.24.16`).
   - Select **Next: Node Spec**.

3. **Step 2: Node Spec.**
   - **Keypair** (required): pick an SSH key pair for node access. The wizard exposes an inline **Create Keypair** button when you don't already have one. The field pre-fills with your most recently used key when at least one key pair exists in the project.
   - **Number of Master Nodes**: defaults to `1`. Type `3` for a production HA control plane. A single-master cluster cannot add masters later.
   - **Flavor of Master Nodes**: select the master flavor (for example, `c2a.xlarge` for 4 vCPU and 8 GB RAM). Overrides the template default.
   - **Number of Nodes**: defaults to `1`. Use `1` for development; size up for workload capacity.
   - **Flavor of Nodes**: select the worker flavor (for example, `c2a.large` for 2 vCPU and 4 GB RAM). Overrides the template default.
   - Select **Next: Network Setting**.

4. **Step 3: Network Setting.**
   - **Enable Load Balancer**: toggle on to create a single endpoint across the master nodes. The checkbox label reads **Enabled Load Balancer for Master Nodes**. The `master_lb_floating_ip_enabled` label on Step 5 controls whether the endpoint gets a floating IP. Default: off.
   - **Enabled Network**: toggle on to auto-provision a dedicated cluster network. The checkbox label reads **Create New Network**. Default: on. Uncheck only when you intend to attach the cluster to a pre-existing network outside this wizard.
   - Select **Next: Management**.

5. **Step 4: Management.**
   - **Auto Healing**: enable to have Magnum restart unhealthy worker nodes. The helper text under the checkbox reads `Automatically repair unhealhty nodes` (the Cloud Console string has a typo: `unhealhty`. The behavior is correct).
   - **Auto Scaling**: enable to adjust the worker count based on demand. Set min/max bounds when enabled.
   - **Timeout(Minute)**: maximum minutes to wait for cluster creation. Default `60`.
   - Select **Next: Additional Labels**.

6. **Step 5: Additional Labels.** This step shows the labels carried by the cluster template you selected on Step 1, with **+ Add Label** to append more and a delete control on each row. The cluster wizard does not inject labels of its own; whatever the template carries is what you see here. A template created without labels (the Console template wizard ships empty by default) renders this step empty.

   The platform requires `boot_volume_size=40` for cluster creation to succeed because every Quake AI flavor family is zero-disk. Confirm `boot_volume_size=40` is present (either inherited from the template or added here via **+ Add Label**) before selecting **Confirm**. Without it, the cluster transitions to **CREATE_FAILED** within about 60 seconds with `Forbidden: Only volume-backed servers are allowed for flavors with zero disk (HTTP 403)`. See [Sizing and quota](/docs/kubernetes/how-to/create-cluster-template#sizing-and-quota) on the template page for the canonical label set the platform-provided templates carry.

   Add `floating_ip_enabled = false` via **+ Add Label** to suppress node-level floating IPs (combine with the template's `master_lb_floating_ip_enabled = true` for a 1-FIP cluster).

7. Select **Confirm**.

The cluster appears in the **Clusters** list with status **CREATE_IN_PROGRESS**. Creation typically takes 5 to 15 minutes depending on cluster size. The status changes to **CREATE_COMPLETE** when the cluster is ready.

</Method>
<Method label="CLI">

Create a cluster from an existing template:

```bash
openstack coe cluster create MY_CLUSTER_NAME \
  --cluster-template MY_TEMPLATE_NAME \
  --keypair MY_KEYPAIR \
  --master-count 1 \
  --master-flavor c2a.xlarge \
  --node-count 1 \
  --flavor c2a.large \
  --timeout 60
```

The example sizes a minimum-viable 1-master + 1-worker cluster that fits a fresh default-tier project (6000 of 8000 `compute_units`). For production HA, pass `--master-count 3` and verify the project's `compute_units` budget covers `3 * (master flavor compute_units) + (worker_count * worker flavor compute_units)` before applying. See [Sizing and quota](/docs/kubernetes/how-to/create-cluster-template#sizing-and-quota) on the template page for the flavor-to-`compute_units` table.

Key parameters:

| Flag | Description |
|---|---|
| `--cluster-template` | Name or ID of the cluster template. Run `openstack coe cluster template list` to see available templates. |
| `--keypair` | SSH key pair name for node access. |
| `--master-count` | Number of master nodes. The wizard defaults to `1`; pass `3` for a production HA control plane. A single-master cluster cannot add masters later. |
| `--node-count` | Number of worker nodes. |
| `--master-flavor` | Override the template's flavor for the master nodes. |
| `--flavor` | Override the template's flavor for the worker nodes. |
| `--timeout` | Maximum minutes to wait for creation (default: 60). |
| `--labels` | Additional labels as comma-separated key-value pairs. Magnum merges these with the template's labels at creation time. Platform templates typically carry `boot_volume_size` and `master_lb_floating_ip_enabled`; pass them here to override per cluster. |

The command returns immediately. The cluster provisions in the background.

Check creation status:

```bash
openstack coe cluster show MY_CLUSTER_NAME -f value -c status
```

Wait until the status shows `CREATE_COMPLETE`.

</Method>
</MethodTabs>

## Verify the result

<MethodTabs>
<Method label="Console">

1. Go to **Kubernetes** > **Clusters**.
2. Confirm the cluster status is **CREATE_COMPLETE** and the health status is healthy.
3. Open the cluster detail page by navigating directly to its URL: `/container-infra/clusters/detail/<cluster-uuid>`. The cluster name in the **ID/Name** column is not a clickable link, and the gear-menu dropdown does not include a "View detail" item; direct URL navigation is the only path to the detail page today.
4. The detail page surfaces fields in seven named sections:

   | Section | What you see |
   |---|---|
   | Top-level | `Name`, `Status`, `Status Reason`, `Health Status`, `Created At`, `Updated At` |
   | Cluster Template | `Name`, `COE` |
   | Network | `Fixed Network`, `Fixed Subnet` |
   | Nodes | `Master Node Flavor`, `Number of Master Nodes`, `Node Flavor`, `Number of Nodes`, `API Address`, `Master Node Addresses`, `Node Addresses` |
   | Additional Labels | `Labels` (the resolved label set: template labels plus any overrides you added on Step 5) |
   | Miscellaneous | `Discovery URL`, `Timeout(Minute)`, `Keypair`, `Docker Volume Size (GiB)`, `COE Version`, `Container Version` |
   | Stack | `Stack ID`, `Stack Faults` (the underlying Heat orchestration handle) |

5. Verify the provisioned infrastructure:
   - **Compute** > **Instances**: master and worker VMs appear with cluster-prefixed names.
   - **Network** > **Networks**: a dedicated cluster network appears (when **Create New Network** was on for Step 3).
   - Run `openstack coe cluster show MY_CLUSTER_NAME` and confirm the API address is populated when **Enable Load Balancer** was on in Step 3.


**False deletion warning on flavor fields.** The cluster detail page renders both flavor fields as `<flavor-name> (The resource has been deleted)` even for healthy clusters with active flavors. The flavors are not deleted; the warning is an upstream Magnum and Cloud Console rendering quirk. Confirm by checking **Compute** > **Instances**, which shows the flavor as active.


</Method>
<Method label="CLI">

View cluster details:

```bash
openstack coe cluster show MY_CLUSTER_NAME
```

Confirm `status` is `CREATE_COMPLETE` and `api_address` has a reachable endpoint.

To access the cluster with `kubectl`, retrieve the kubeconfig:

```bash
openstack coe cluster config MY_CLUSTER_NAME
```

This outputs an `export KUBECONFIG=...` command. Run it, then verify connectivity:

```bash
kubectl get nodes
```

Expected output for a 1-master + 1-worker cluster:

```text
NAME                                       STATUS   ROLES    AGE   VERSION
my-cluster-<random-token>-master-0         Ready    master   10m   v1.24.16
my-cluster-<random-token>-node-0           Ready    <none>   8m    v1.24.16
```

Magnum injects a 12-character random token between the cluster name and the role suffix to keep node hostnames unique across cluster re-creates. Magnum names worker nodes `node-<index>`; the kubectl `ROLES` column shows `<none>` for workers.

If `kubectl` cannot reach the API endpoint and the cluster reports `master is not Ready`, check the project's floating IP budget. The LB VIP FIP allocates as part of `tofu apply` or `openstack coe cluster create`; FIP-quota exhaustion can leave the cluster in a half-up state without surfacing the cause in `status_reason`. Run `openstack quota show` for the current project allocation.

</Method>
</MethodTabs>

## Next steps

- [How to manage a Kubernetes cluster](/docs/kubernetes/how-to/manage-cluster): scale nodes, view status, and delete clusters
- [Kubernetes on Quake AI](/docs/kubernetes/concepts/kubernetes): architecture and resource planning
- [Clusters Console](/reference/kubernetes/console/clusters): monitor clusters in the console
- [Kubernetes CLI reference](/reference/kubernetes/cli)
