How to auto-scale a VM tier with Heat
Coming from another cloud?
▸AWS·Stacks
Stacks
- Quake AI offers Heat (legacy OpenStack orchestration) and recommends OpenTofu for new IaC work; EKS uses declarative console/CLI.
- EKS Auto Mode automates data plane; Quake AI uses OpenTofu modules for infrastructure provisioning.
- CloudFormation stacks are managed via AWS-specific REST API (e.g., cloudformation.us-east-1.amazonaws.com) requiring AWS SigV4 auth, while Heat (legacy) uses OpenStack Identity API v3 (keystoneauth) and OpenTofu uses the OpenStack provider with application credentials. Migrators notice different endpoint discovery and auth flows.
- Stacks support StackSets for cross-region/account deployment; Heat has no equivalent. OpenTofu workspaces offer a different multi-environment pattern.
▸Azure·Resource Manager deployment (deployment)
Azure Resource Manager deployment (deployment)
- Authoring format differs: ARM templates are JSON documents for Azure Resource Manager deployments, whereas Quake AI offers Heat (legacy, HOT/YAML format) and recommends OpenTofu (HCL) for new infrastructure automation.
- Deployment target/scope differs: ARM templates can deploy at multiple scopes (resource group, subscription, management group, tenant) via Azure Resource Manager's management layer, while Heat orchestrates within a single OpenStack project/tenant. OpenTofu can target multiple projects via provider configuration.
- Change-preview behavior differs: ARM supports a built-in “what-if” style preview of changes via Resource Manager deployments, while OpenTofu provides `tofu plan` for change previews. Heat relies on stack update events (legacy workflow).
- Resource coverage lifecycle differs: ARM templates support Azure resource types exposed by Azure resource providers; OpenTofu covers OpenStack resource types via the OpenStack provider, and Heat (legacy) supports a narrower set via built-in resource plugins.
▸Google Cloud·Deployment Manager
Deployment Manager
- Uses YAML configurations with optional Jinja2/Python templates expanded server-side; Quake AI offers Heat (legacy, HOT/YAML) and recommends OpenTofu (HCL) for new work.
- Resource types named as service.v1.resource (e.g., compute.v1.instance) tied to GCP APIs, vs OpenStack prefixed types like OS::Nova::Server.
- Single integrated GCP service (no separate API/engine processes); Heat has a multi-service architecture (heat-api, heat-engine) but is legacy. OpenTofu runs client-side with no server component.
- No additional service fee; bills only for deployed GCP resources, same as OpenStack resources billing via provider.
How to auto-scale a VM tier with Heat
Run a private worker tier that grows and shrinks by combining an OS::Heat::AutoScalingGroup with scale-up and scale-down policies. Each policy exposes a signal webhook that your monitoring or queue system can call.
Prerequisites
- CLIOpenStack CLI installed and authenticated (
clouds.yamloropenrcsourced) - ConsoleLogged in to the Quake AI console
Windows: CLI examples use bash. Set up a Linux CLI environment on Windows before proceeding.
- An unrestricted application credential. Heat creates a Keystone trust to drive convergence.
- An SSH key pair in your project.
- Familiarity with HOT YAML.
- A private service, queue, or job source that group members can reach.
What the stack creates#
The stack provisions:
- A private network, subnet, and router.
- An
OS::Heat::AutoScalingGroupof volume-backed worker instances. - Scale-up and scale-down
OS::Heat::ScalingPolicyresources. - Signal webhooks that add or remove one worker.
The group members stay on private addresses. The project router provides outbound access for package installation and job retrieval.
Define the worker template#
Create autoscale-worker.yaml for one group member:
heat_template_version: 2021-04-16
description: One volume-backed worker for a Heat autoscaling group
parameters:
flavor:
type: string
image:
type: string
key_name:
type: string
network:
type: string
security_group:
type: string
resources:
port:
type: OS::Neutron::Port
properties:
network: { get_param: network }
security_groups:
- { get_param: security_group }
server:
type: OS::Nova::Server
properties:
flavor: { get_param: flavor }
key_name: { get_param: key_name }
networks:
- port: { get_resource: port }
block_device_mapping_v2:
- boot_index: 0
image: { get_param: image }
volume_size: 20
delete_on_termination: true
user_data_format: RAW
user_data: |
#!/bin/bash
apt-get update
apt-get install -y docker.io
systemctl enable --now docker
docker run -d --restart unless-stopped \
--name worker \
YOUR_REGISTRY/YOUR_WORKER_IMAGE:YOUR_TAG
outputs:
private_address:
value: { get_attr: [server, first_address] }Replace the image reference with your worker image. Pass queue credentials through your deployment system or instance bootstrap secret source. Do not embed credentials in the template.
Define the autoscaling group#
Create autoscale-workers.yaml in the same directory:
heat_template_version: 2021-04-16
description: Private worker tier with Heat scaling policies
parameters:
flavor:
type: string
default: s1a.small
image:
type: string
default: Ubuntu-24.04
key_name:
type: string
external_network:
type: string
default: PublicStatic
private_cidr:
type: string
default: 192.168.70.0/24
min_size:
type: number
default: 1
max_size:
type: number
default: 6
resources:
network:
type: OS::Neutron::Net
properties:
name: autoscale-workers-net
subnet:
type: OS::Neutron::Subnet
properties:
network: { get_resource: network }
cidr: { get_param: private_cidr }
ip_version: 4
enable_dhcp: true
dns_nameservers: [8.8.8.8, 8.8.4.4]
router:
type: OS::Neutron::Router
properties:
external_gateway_info:
network: { get_param: external_network }
router_interface:
type: OS::Neutron::RouterInterface
properties:
router: { get_resource: router }
subnet: { get_resource: subnet }
worker_security_group:
type: OS::Neutron::SecurityGroup
properties:
description: Egress-only worker group
worker_group:
type: OS::Heat::AutoScalingGroup
depends_on: router_interface
properties:
min_size: { get_param: min_size }
max_size: { get_param: max_size }
desired_capacity: { get_param: min_size }
resource:
type: autoscale-worker.yaml
properties:
flavor: { get_param: flavor }
image: { get_param: image }
key_name: { get_param: key_name }
network: { get_resource: network }
security_group: { get_resource: worker_security_group }
scale_up:
type: OS::Heat::ScalingPolicy
properties:
auto_scaling_group_id: { get_resource: worker_group }
adjustment_type: change_in_capacity
scaling_adjustment: 1
cooldown: 60
scale_down:
type: OS::Heat::ScalingPolicy
properties:
auto_scaling_group_id: { get_resource: worker_group }
adjustment_type: change_in_capacity
scaling_adjustment: -1
cooldown: 60
outputs:
scale_up_url:
value: { get_attr: [scale_up, alarm_url] }
scale_down_url:
value: { get_attr: [scale_down, alarm_url] }Quake AI flavors report zero root disk, so the member template boots each instance from a volume through block_device_mapping_v2.
Deploy the stack#
Deploy from the directory that contains both YAML files:
openstack stack create \
--template autoscale-workers.yaml \
--parameter key_name=YOUR_KEYPAIR \
--parameter external_network=PublicStatic \
autoscale-workersWatch the stack reach CREATE_COMPLETE:
openstack stack show autoscale-workers \
-c stack_status \
-c stack_status_reasonVerify the group#
Trigger a scale event#
Add one worker:
openstack stack resource signal autoscale-workers scale_upAfter the 60-second cool-down period, list the nested server resources again. Signal scale_down to remove one worker:
openstack stack resource signal autoscale-workers scale_downHeat keeps the group between min_size and max_size.
Connect an external trigger#
Read the policy URLs:
openstack stack output show autoscale-workers scale_up_url \
-c output_value \
-f value
openstack stack output show autoscale-workers scale_down_url \
-c output_value \
-f valueConfigure your queue monitor, CI system, or self-managed monitoring stack to call the appropriate URL when capacity should change. The Monitoring Stack template provides Prometheus and Grafana when you need a monitoring host.
Treat policy URLs as credentials. Store them in the triggering system's secret store and rotate them by replacing the policy resource if a URL leaks.
Tune scaling behavior#
- Change
min_sizeandmax_sizeto set group bounds. - Change
scaling_adjustmentto add or remove several instances per signal. - Increase
cooldownwhen workers need more time to boot and begin processing. - Use queue depth, pending job age, or sustained CPU use as the trigger signal.
- Add graceful shutdown logic so a worker finishes or returns its current job before scale-down removes it.
See also#
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.