# How to provision a Quake AI instance with Ansible

Source: https://docs.quake.ai/docs/automation/how-to/ansible-provision-instance
Markdown: https://docs.quake.ai/docs/automation/how-to/ansible-provision-instance.md

---

# How to provision a Quake AI instance with Ansible

This guide creates a complete Quake AI stack (network, subnet, router, security group, keypair, instance, floating IP) using only the `openstack.cloud` Ansible collection. No OpenTofu, no separate provisioning tool. The result is a single playbook that brings up infrastructure and a paired teardown playbook that removes it.

This pattern fits small deployments, proof of concept work, and teams that are already fully Ansible-native and want to avoid adopting a second tool. It does not fit production fleets that need state-managed plan-and-apply workflows.



Ansible has no state file, no plan step, and no built-in drift detection. Running the same provisioning playbook twice creates resources idempotently, but Ansible cannot tell you what would change before you run it, and it cannot detect resources that drifted out of band. For teams managing multiple environments or large infrastructure, [OpenTofu](/docs/automation/concepts/iac-comparison) remains the better fit for provisioning. The trade-off rationale lives in [Ansible on Quake AI](/docs/automation/concepts/ansible).



## Prerequisites

- [Ansible setup](/docs/automation/how-to/getting-started-ansible) complete: `ansible`, `openstacksdk >= 1.0.0`, the `openstack.cloud` collection installed
- A named cloud (`rumble`) configured in `~/.config/openstack/clouds.yaml`. The provisioning modules use the same authentication chain as the dynamic inventory plugin and the configuration playbooks. Re-document of that flow lives in the [getting-started guide](/docs/automation/how-to/getting-started-ansible#configure-authentication).
- A local SSH public key at `~/.ssh/id_ed25519.pub`

## When to use this pattern

Use the all-in-one Ansible path when:

- The whole fleet fits in a single playbook and a single team
- The team has Ansible expertise and no OpenTofu expertise
- Day 0 (provision) and Day 1 (configure) live in the same repo and never run independently

Reach for [OpenTofu plus Ansible](/docs/automation/how-to/ansible-opentofu-workflow) instead when:

- Multiple environments share infrastructure definitions (dev, staging, production)
- A provisioning change should be reviewed in a plan before it applies
- The infrastructure footprint is large enough that drift matters

## Project layout

Create the project structure:

```bash
mkdir ansible-provision-demo && cd ansible-provision-demo
mkdir -p playbooks/group_vars
```

Add `playbooks/group_vars/all.yml` beside the playbooks so Ansible loads the vars when you run `playbooks/provision.yml`:

```yaml
cloud_name: rumble
project_name: ansible-demo

network_name: ansible-demo-net
subnet_name: ansible-demo-subnet
subnet_cidr: 10.20.0.0/24
router_name: ansible-demo-router
# us-east-1 external network for floating-IP allocation. Discover with
# `openstack network list --external` if you are on a different region.
external_network: PublicStatic

security_group_name: ansible-demo-web-sg
keypair_name: ansible-demo-key
keypair_public_key_file: "{{ lookup('env', 'HOME') }}/.ssh/id_ed25519.pub"

server_name: ansible-demo-web-01
server_image: Ubuntu-24.04
# Quake AI flavors ship with disk: 0, so the server task below boots
# from a Cinder volume. Discover live flavor names with:
#   openstack flavor list -f value -c Name -c Disk
server_flavor: m2a.large
server_boot_volume_size: 20
```

The values match the platform defaults documented in [Ansible on Quake AI](/docs/automation/concepts/ansible). Adjust `subnet_cidr`, `server_flavor`, `server_image`, or `server_boot_volume_size` to fit your project. Every Quake AI flavor family currently advertises `disk: 0`, which is why the provisioning task below sets `boot_from_volume: true`.

## Build the provisioning playbook

Create `playbooks/provision.yml`:

```yaml
---
- name: Provision Quake AI stack
  hosts: localhost
  gather_facts: false
  module_defaults:
    group/openstack.cloud.openstack:
      cloud: "{{ cloud_name }}"

  tasks:
    - name: Create network
      openstack.cloud.network:
        state: present
        name: "{{ network_name }}"
      register: net

    - name: Create subnet
      openstack.cloud.subnet:
        state: present
        name: "{{ subnet_name }}"
        network_name: "{{ network_name }}"
        cidr: "{{ subnet_cidr }}"
        ip_version: 4
        dns_nameservers:
          - 1.1.1.1
          - 1.0.0.1
      register: subnet

    - name: Create router
      openstack.cloud.router:
        state: present
        name: "{{ router_name }}"
        network: "{{ external_network }}"
        interfaces:
          - "{{ subnet_name }}"

    - name: Create security group
      openstack.cloud.security_group:
        state: present
        name: "{{ security_group_name }}"
        description: Web ingress for the ansible-demo project
      register: sg

    - name: Allow SSH
      openstack.cloud.security_group_rule:
        state: present
        security_group: "{{ security_group_name }}"
        protocol: tcp
        port_range_min: 22
        port_range_max: 22
        remote_ip_prefix: 0.0.0.0/0

    - name: Allow HTTP
      openstack.cloud.security_group_rule:
        state: present
        security_group: "{{ security_group_name }}"
        protocol: tcp
        port_range_min: 80
        port_range_max: 80
        remote_ip_prefix: 0.0.0.0/0

    - name: Upload SSH keypair
      openstack.cloud.keypair:
        state: present
        name: "{{ keypair_name }}"
        public_key_file: "{{ keypair_public_key_file }}"

    - name: Create server
      openstack.cloud.server:
        state: present
        name: "{{ server_name }}"
        image: "{{ server_image }}"
        flavor: "{{ server_flavor }}"
        boot_from_volume: true
        volume_size: "{{ server_boot_volume_size }}"
        terminate_volume: true
        key_name: "{{ keypair_name }}"
        security_groups:
          - "{{ security_group_name }}"
        network: "{{ network_name }}"
        auto_ip: false
        wait: true
        meta:
          environment: demo
          role: web
      register: server

    - name: Allocate floating IP
      openstack.cloud.floating_ip:
        state: present
        server: "{{ server_name }}"
        network: "{{ external_network }}"
        wait: true
      register: fip

    - name: Show floating IP
      ansible.builtin.debug:
        msg: "Server {{ server_name }} reachable at {{ fip.floating_ip.floating_ip_address }}"
```

The `module_defaults` block sets `cloud: rumble` for every `openstack.cloud.*` task in this play. Without it, every module needs an explicit `cloud:` parameter.

The `auto_ip: false` plus separate `openstack.cloud.floating_ip` task is the recommended pattern: `auto_ip: true` works but couples the server task to floating IP allocation. Splitting the steps makes failures easier to diagnose and reruns cleaner.

The `boot_from_volume: true` plus `volume_size:` plus `terminate_volume: true` triplet is required because every Quake AI flavor family ships with `disk: 0`. Without those three keys, `openstack.cloud.server` fails with `Only volume-backed servers are allowed for flavors with zero disk`. The Cinder volume that backs the boot disk is sized by `server_boot_volume_size` (in GiB) and is released alongside the server because of `terminate_volume: true`. To keep the boot disk after the server is deleted (for forensic capture or reuse), set `terminate_volume: false`.

## Run the playbook

Provision the whole stack:

```bash
ansible-playbook playbooks/provision.yml
```

A first run prints `changed=10` (or similar) as each resource is created. Idempotent re-runs of the same playbook print `changed=0` because every module reconciles to the desired state.

Verify the instance is reachable:

```bash
ssh -i ~/.ssh/id_ed25519 ubuntu@<floating_ip_from_debug_task>
```

The floating IP is also visible in the OpenStack CLI:

```bash
openstack --os-cloud quakeai server list
openstack --os-cloud quakeai floating ip list
```

## Author the teardown playbook

Ansible has no `destroy` equivalent of `tofu destroy`. Removal happens by running tasks with `state: absent` against the same resources, in reverse dependency order.

Create `playbooks/teardown.yml`:

```yaml
---
- name: Tear down Quake AI stack
  hosts: localhost
  gather_facts: false
  module_defaults:
    group/openstack.cloud.openstack:
      cloud: "{{ cloud_name }}"

  tasks:
    - name: Look up floating IP for server
      openstack.cloud.floating_ip_info:
        server: "{{ server_name }}"
      register: server_fips

    - name: Remove floating IP
      openstack.cloud.floating_ip:
        state: absent
        server: "{{ server_name }}"
        floating_ip_address: "{{ server_fips.floating_ips[0].floating_ip_address }}"
        network: "{{ external_network }}"
      when: server_fips.floating_ips | length > 0

    - name: Delete server
      openstack.cloud.server:
        state: absent
        name: "{{ server_name }}"
        wait: true

    - name: Delete keypair
      openstack.cloud.keypair:
        state: absent
        name: "{{ keypair_name }}"

    - name: Delete security group
      openstack.cloud.security_group:
        state: absent
        name: "{{ security_group_name }}"

    - name: Detach router from subnet
      openstack.cloud.router:
        state: absent
        name: "{{ router_name }}"

    - name: Delete subnet
      openstack.cloud.subnet:
        state: absent
        name: "{{ subnet_name }}"

    - name: Delete network
      openstack.cloud.network:
        state: absent
        name: "{{ network_name }}"
```

Run the teardown:

```bash
ansible-playbook playbooks/teardown.yml
```

The order matters. Floating IPs release before the server they point at, the server releases before its keypair and security group, the router detaches before its subnet, and the subnet deletes before the network.

On `openstack.cloud` 2.5.0, `state: absent` on `openstack.cloud.floating_ip` requires `server`, `floating_ip_address`, and `network` together. The lookup task resolves the address from the live server before the removal task runs.

## Modules referenced

This playbook touches the core compute and network modules in the `openstack.cloud` collection:

| Module | Role in the playbook |
|---|---|
| `openstack.cloud.network` | Create the project network |
| `openstack.cloud.subnet` | Allocate a subnet inside the network |
| `openstack.cloud.router` | Attach the subnet to the external network |
| `openstack.cloud.security_group` | Define the per-project ingress group |
| `openstack.cloud.security_group_rule` | Allow SSH and HTTP into the group |
| `openstack.cloud.keypair` | Upload the public key the server will trust |
| `openstack.cloud.server` | Create the instance and tag it with metadata |
| `openstack.cloud.floating_ip` | Allocate and associate the public address |

For block storage, swap or extend the playbook with `openstack.cloud.volume` and `openstack.cloud.server_volume`.

## Configure the host

Provisioning the instance is half the workflow. Once `provision.yml` finishes, the floating IP exists but nothing is configured on the host. The clean handoff is to run a separate configuration play that targets the instance over SSH.

Tell Ansible which inventory to use. The simplest path is the [dynamic inventory plugin](/docs/automation/how-to/ansible-dynamic-inventory), which picks up the new instance via the `role: web` and `environment: demo` metadata you set during provisioning. Then the configure play looks like:

```yaml
- name: Configure web host
  hosts: role_web
  become: true
  tasks:
    - ansible.builtin.wait_for_connection: { timeout: 300 }
    - ansible.builtin.apt: { update_cache: true }
    - ansible.builtin.apt: { name: nginx, state: present }
```

For larger projects, the [OpenTofu plus Ansible workflow](/docs/automation/how-to/ansible-opentofu-workflow) covers the same handoff with state-managed provisioning.

## See also

- [Ansible on Quake AI](/docs/automation/concepts/ansible)
- [How to get started with Ansible on Quake AI](/docs/automation/how-to/getting-started-ansible)
- [How to use Ansible dynamic inventory on Quake AI](/docs/automation/how-to/ansible-dynamic-inventory)
- [How to use Ansible with OpenTofu on Quake AI](/docs/automation/how-to/ansible-opentofu-workflow)
- [IaC on Quake AI: OpenTofu and Terraform](/docs/automation/concepts/iac-comparison)
- [Generate app credentials](/docs/tools/generate-app-credentials)
