# How to auto-scale a VM tier with Heat

Source: https://docs.quake.ai/docs/automation/how-to/autoscale-with-heat
Markdown: https://docs.quake.ai/docs/automation/how-to/autoscale-with-heat.md

---

# How to auto-scale a VM tier with Heat

Run a private worker tier that grows and shrinks by combining an `OS::Heat::AutoScalingGroup` with scale-up and scale-down policies. Each policy exposes a signal webhook that your monitoring or queue system can call.



This pattern suits queue consumers, build workers, render workers, and other instances that do not need a public listener. For a public application entry point, run a fixed [edge reverse proxy](/resources/iac-templates/edge-reverse-proxy) or [API gateway](/resources/iac-templates/api-gateway) and update its private backend discovery when the group changes.



<PrerequisiteBlock methods={["cli", "console"]}>

- An unrestricted application credential. Heat creates a Keystone trust to drive convergence.
- An SSH key pair in your project.
- Familiarity with HOT YAML.
- A private service, queue, or job source that group members can reach.

</PrerequisiteBlock>

## What the stack creates

The stack provisions:

- A private network, subnet, and router.
- An `OS::Heat::AutoScalingGroup` of volume-backed worker instances.
- Scale-up and scale-down `OS::Heat::ScalingPolicy` resources.
- Signal webhooks that add or remove one worker.

The group members stay on private addresses. The project router provides outbound access for package installation and job retrieval.

## Define the worker template

Create `autoscale-worker.yaml` for one group member:

```yaml
heat_template_version: 2021-04-16

description: One volume-backed worker for a Heat autoscaling group

parameters:
  flavor:
    type: string
  image:
    type: string
  key_name:
    type: string
  network:
    type: string
  security_group:
    type: string

resources:
  port:
    type: OS::Neutron::Port
    properties:
      network: { get_param: network }
      security_groups:
        - { get_param: security_group }

  server:
    type: OS::Nova::Server
    properties:
      flavor: { get_param: flavor }
      key_name: { get_param: key_name }
      networks:
        - port: { get_resource: port }
      block_device_mapping_v2:
        - boot_index: 0
          image: { get_param: image }
          volume_size: 20
          delete_on_termination: true
      user_data_format: RAW
      user_data: |
        #!/bin/bash
        apt-get update
        apt-get install -y docker.io
        systemctl enable --now docker
        docker run -d --restart unless-stopped \
          --name worker \
          YOUR_REGISTRY/YOUR_WORKER_IMAGE:YOUR_TAG

outputs:
  private_address:
    value: { get_attr: [server, first_address] }
```

Replace the image reference with your worker image. Pass queue credentials through your deployment system or instance bootstrap secret source. Do not embed credentials in the template.

## Define the autoscaling group

Create `autoscale-workers.yaml` in the same directory:

```yaml
heat_template_version: 2021-04-16

description: Private worker tier with Heat scaling policies

parameters:
  flavor:
    type: string
    default: s1a.small
  image:
    type: string
    default: Ubuntu-24.04
  key_name:
    type: string
  external_network:
    type: string
    default: PublicStatic
  private_cidr:
    type: string
    default: 192.168.70.0/24
  min_size:
    type: number
    default: 1
  max_size:
    type: number
    default: 6

resources:
  network:
    type: OS::Neutron::Net
    properties:
      name: autoscale-workers-net

  subnet:
    type: OS::Neutron::Subnet
    properties:
      network: { get_resource: network }
      cidr: { get_param: private_cidr }
      ip_version: 4
      enable_dhcp: true
      dns_nameservers: [8.8.8.8, 8.8.4.4]

  router:
    type: OS::Neutron::Router
    properties:
      external_gateway_info:
        network: { get_param: external_network }

  router_interface:
    type: OS::Neutron::RouterInterface
    properties:
      router: { get_resource: router }
      subnet: { get_resource: subnet }

  worker_security_group:
    type: OS::Neutron::SecurityGroup
    properties:
      description: Egress-only worker group

  worker_group:
    type: OS::Heat::AutoScalingGroup
    depends_on: router_interface
    properties:
      min_size: { get_param: min_size }
      max_size: { get_param: max_size }
      desired_capacity: { get_param: min_size }
      resource:
        type: autoscale-worker.yaml
        properties:
          flavor: { get_param: flavor }
          image: { get_param: image }
          key_name: { get_param: key_name }
          network: { get_resource: network }
          security_group: { get_resource: worker_security_group }

  scale_up:
    type: OS::Heat::ScalingPolicy
    properties:
      auto_scaling_group_id: { get_resource: worker_group }
      adjustment_type: change_in_capacity
      scaling_adjustment: 1
      cooldown: 60

  scale_down:
    type: OS::Heat::ScalingPolicy
    properties:
      auto_scaling_group_id: { get_resource: worker_group }
      adjustment_type: change_in_capacity
      scaling_adjustment: -1
      cooldown: 60

outputs:
  scale_up_url:
    value: { get_attr: [scale_up, alarm_url] }
  scale_down_url:
    value: { get_attr: [scale_down, alarm_url] }
```

Quake AI flavors report zero root disk, so the member template boots each instance from a volume through `block_device_mapping_v2`.

## Deploy the stack

Deploy from the directory that contains both YAML files:

```bash
openstack stack create \
  --template autoscale-workers.yaml \
  --parameter key_name=YOUR_KEYPAIR \
  --parameter external_network=PublicStatic \
  autoscale-workers
```

Watch the stack reach `CREATE_COMPLETE`:

```bash
openstack stack show autoscale-workers \
  -c stack_status \
  -c stack_status_reason
```

## Verify the group

<MethodTabs>
<Method label="Console">

1. Select **Automation** > **Heat Stacks**.
2. Open `autoscale-workers`.
3. On **Stack Resources**, find `worker_group`.
4. Confirm the group and its nested server resources show `Create Complete`.

</Method>
<Method label="CLI">

List the worker instances through the nested stack:

```bash
openstack stack resource list autoscale-workers -n 2 \
  --filter type=OS::Nova::Server \
  -c resource_name \
  -c physical_resource_id
```

</Method>
</MethodTabs>

## Trigger a scale event

Add one worker:

```bash
openstack stack resource signal autoscale-workers scale_up
```

After the 60-second cool-down period, list the nested server resources again. Signal `scale_down` to remove one worker:

```bash
openstack stack resource signal autoscale-workers scale_down
```

Heat keeps the group between `min_size` and `max_size`.

## Connect an external trigger

Read the policy URLs:

```bash
openstack stack output show autoscale-workers scale_up_url \
  -c output_value \
  -f value

openstack stack output show autoscale-workers scale_down_url \
  -c output_value \
  -f value
```

Configure your queue monitor, CI system, or self-managed monitoring stack to call the appropriate URL when capacity should change. The [Monitoring Stack template](/resources/iac-templates/monitoring-stack) provides Prometheus and Grafana when you need a monitoring host.

Treat policy URLs as credentials. Store them in the triggering system's secret store and rotate them by replacing the policy resource if a URL leaks.

## Tune scaling behavior

- Change `min_size` and `max_size` to set group bounds.
- Change `scaling_adjustment` to add or remove several instances per signal.
- Increase `cooldown` when workers need more time to boot and begin processing.
- Use queue depth, pending job age, or sustained CPU use as the trigger signal.
- Add graceful shutdown logic so a worker finishes or returns its current job before scale-down removes it.

## See also

- [How to create a Heat stack](/docs/automation/how-to/create-heat-stack)
- [Heat and OpenTofu compared](/docs/automation/concepts/iac-comparison)
- [Heat Simple Stack template](/resources/iac-templates/heat-simple-stack)
- [Monitoring Stack template](/resources/iac-templates/monitoring-stack)
- [Edge Reverse Proxy template](/resources/iac-templates/edge-reverse-proxy)
