# Kubernetes

Source: https://docs.quake.ai/docs/kubernetes
Markdown: https://docs.quake.ai/docs/kubernetes.md
> Provision Kubernetes clusters on Quake AI from reusable templates: the platform creates master nodes, worker nodes, and networking with Magnum.

---

# Kubernetes

The Kubernetes service automates cluster provisioning on Quake AI. You define the node sizes, network driver, and Kubernetes version through a reusable template, and the platform provisions master nodes, worker nodes, and the networking between them using existing Compute, Network, and Storage infrastructure.

[OpenStack Magnum](https://docs.openstack.org/magnum/latest/) backs Kubernetes. For details on how Quake AI implements OpenStack, see [How Quake AI uses OpenStack](/resources/migration/openstack).



The Kubernetes service automates cluster provisioning on Quake AI: it creates the VMs, networking, and stable control-plane endpoints you need to run Kubernetes from a reusable template. Master nodes run in your project as Compute instances, which gives you full visibility into control plane resources and lets you tune HA with anti-affinity server groups, etcd backups, and your own upgrade schedule.

Managed Kubernetes offerings from EKS, GKE, and AKS hide the control plane and include a platform SLA. Quake AI targets teams comfortable operating Kubernetes who want automated infrastructure provisioning without a fully managed control plane contract.



## What you can do

<UseCaseGrid>
<UseCaseCard
  title="Create a Kubernetes cluster"
  description="Deploy a production-ready Kubernetes cluster with master and worker nodes from a cluster template."
  href="/docs/kubernetes/how-to/create-cluster"
  difficulty="intermediate"
  estimatedTime="20 min"
  services={["Kubernetes", "Compute", "Network"]}
/>
<UseCaseCard
  title="Define a cluster template"
  description="Create a reusable blueprint that specifies Kubernetes version, node flavors, network settings, and scaling policies."
  href="/docs/kubernetes/how-to/create-cluster-template"
  difficulty="intermediate"
  estimatedTime="15 min"
  services={["Kubernetes"]}
/>
<UseCaseCard
  title="Scale and manage clusters"
  description="Add or remove worker nodes, update configurations, and manage the lifecycle of running clusters."
  href="/docs/kubernetes/how-to/manage-cluster"
  difficulty="intermediate"
  estimatedTime="10 min"
  services={["Kubernetes", "Compute"]}
/>
</UseCaseGrid>

## How it works

<Figure size="md" caption="Kubernetes architecture: a cluster template provisions a control plane and worker plane on Compute, with Network and Storage attached.">

```d2
template: Cluster Template {
  spec: "Image\nFlavor\nK8s version\nNetwork driver\nVolume driver"
}

cluster: Kubernetes Cluster {
  control: Control Plane {
    api: API server
    sched: Scheduler
    cm: Controller manager
    etcd: etcd
  }
  workers: Worker Plane {
    w1: "Worker node\nkubelet + runtime"
    w2: "Worker node\nkubelet + runtime"
  }
  control.api -> workers.w1
  control.api -> workers.w2
}

infra: Quake AI Infrastructure {
  compute: Compute (VMs)
  net: Network (CNI, LB, FIP)
  store: Storage (volumes)
}

user: Operator {
  shape: person
}

template -> cluster: provisions
cluster.control -> infra.compute: runs on
cluster.workers -> infra.compute: runs on
cluster.workers -> infra.net: pod + service network
cluster.workers -> infra.store: persistent volumes
user -> cluster.control.api: kubectl
```

</Figure>

A **cluster template** defines the blueprint: which Kubernetes version to run, the compute flavor for master and worker nodes, the network driver, volume driver for persistent storage, and optional settings like auto-scaling and TLS. You create the template once and reuse it for multiple clusters.

When you create a **cluster** from a template, the platform provisions the underlying infrastructure: master VMs running the Kubernetes API server, scheduler, and controller manager; worker VMs running the kubelet and container runtime; and the networking that connects them. The cluster integrates with Quake AI's Compute service for VM provisioning, Network service for pod and service networking, and Storage service for persistent volumes.

Once the cluster is running, you interact with it through standard Kubernetes tools. Use `kubectl` to deploy workloads, manage pods, and configure services. The Quake AI console, CLI, and API handle cluster-level operations: scaling the node pool, upgrading Kubernetes versions, and deleting clusters when no longer needed.



Kubernetes clusters require a minimum of two VMs (one master, one worker). Production deployments typically use three master nodes with 8 vCPU / 16 GB RAM each plus additional workers. Verify that your project's vCPU quota supports your planned cluster size before creating.



## Get started

To deploy your first cluster, start with [Create a cluster template](/docs/kubernetes/how-to/create-cluster-template), then [Create a Kubernetes cluster](/docs/kubernetes/how-to/create-cluster).

## Key concepts

<DocsSectionLinks section="kubernetes/concepts" grouping="concepts" primaryOnly />

## Security considerations

Cluster security starts with the **cluster template**: it defines the network driver, TLS settings, and node image. Beyond that, harden workloads with Kubernetes-native controls: RBAC for API access, network policies for pod-to-pod traffic, and secrets management for credentials. The underlying VMs inherit compute and network security from the platform. See [Security](/docs/security) for cross-service security documentation.

## Guides and reference

### Console guides

<ReferenceConsoleLinks service="kubernetes" />

### How-to guides

<DocsSectionLinks section="kubernetes/how-to" />

### CLI reference

<ReferenceServiceLinks service="kubernetes" group="cli" />

### API reference

<ReferenceServiceLinks service="kubernetes" group="api" />

## Related services

Kubernetes clusters run on [Compute](/docs/compute) instances and use [Network](/docs/network) for pod connectivity and floating IPs. Kubernetes Services with `type: LoadBalancer` receive public endpoints through the cluster's cloud controller. [Storage](/docs/platform#storage) block volumes back the persistent volumes attached to Kubernetes pods. For infrastructure-as-code cluster provisioning, see [Automation](/docs/automation).
