Skip to content

Deploy a bursty-GPU control plane with a queue and workers

Deployment

Coming from another cloud?

▸AWS·SQS

This Quake AI feature maps to AWS’s SQS.

▸Google Cloud·Pubsub

This Quake AI feature maps to Google Cloud’s Pubsub.

Deploy a bursty-GPU control plane with a queue and workers

Stand up a bursty-GPU control plane on Quake AI CPU compute using the validated OpenTofu template redis-cache for the job queue. A Redis or Valkey queue holds jobs, a worker process on a CPU instance pops each job, submits it to a third-party GPU or model API, polls until the result is ready, and writes the output to Object Storage. The queue and the worker run on instances you own and keep running cheaply; the heavy GPU compute runs on whichever provider you point the worker at.

This is the brain-on-CPU, muscle-elsewhere shape from AI inference and RAG pipelines, made concrete as a job pipeline. The durable parts, the queue, the worker loop, the retry policy, the credentials, and the result store, live on the control plane you operate. The GPU backend is a callout the worker makes, so you can repoint it at a different provider without changing the pipeline around it. OpenClaw runs the same always-on-agent-on-CPU shape as a one-click Console App.

Your app or workstationQuake AI control planeGPU provider (elsewhere)Enqueue jobRedis / Valkey queue(private IP)Worker(CPU instance)Object StorageresultsInference or GPU API BLPOP jobstore resultRPUSHsubmit and poll
Click to zoom
What you will build: an enqueue step pushes jobs to a Redis or Valkey queue on Quake AI; a CPU worker pops each job, submits it to a GPU provider elsewhere, polls for the result, and writes it to Object Storage

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on shared vCPU.

Starting template$13.90/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Cache

s1a.small · 2 shared vCPU, 2 GiB RAM, 0.5 Gbps

$16.50/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

s1a.small

2 shared vCPU, 2 GiB RAM, 0.5 Gbps

$16.50

Compute + RAM rate basis

2 vCPU + 2 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (30 GiB)

30 GiB at $0.08/GiB/mo

$2.40

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Pricing data last validated: . For current rates, check quake.ai/pricing.

The queue is the redis-cache template above. The worker instance, the Object Storage bucket, and the worker code are the custom layer you add on top in the steps below.

Prerequisites#

Before you start, confirm you have:

  • A Quake AI account with application credentials and the OpenStack CLI configured
  • OpenTofu installed locally (installation guide) for the queue step
  • An SSH key pair that already exists in your project; record its name for the key_name variables
  • An API key for one third-party GPU or model provider with an HTTP API, and the base URL of that API
  • Enough project quota for two s1a instances, two volumes, one private network, a router, two security groups, and one floating IP
  • Basic familiarity with Object Storage on Quake AI

This pipeline runs CPU-only on Quake AI. The GPU compute runs on the external provider you choose; Quake AI flavors are AMD EPYC with no GPU option (compute FAQ). The pattern fits batch inference, image and audio generation, and any job you submit to a remote accelerator and collect later.

Step 1: Stand up the queue#

The queue is a Redis or Valkey instance on a private network. Deploy it with the redis-cache template by following Deploy a Redis or Valkey cache with the redis-cache template, then return here. That walkthrough provisions the private network, the security group, the optional persistence volume, and the cache instance, and it leaves the cache reachable only from the private subnet.

When the cache stack is up, capture three values from its tofu output:

bash
CACHE_PRIVATE_IP=$(tofu output -raw private_ip)
echo "$CACHE_PRIVATE_IP"

Record the cache's private IP, the private network name, and the password you set. The worker reaches the cache at redis://:REDIS_PASSWORD@CACHE_PRIVATE_IP:6379/0 over the private network. Leave persistence_enabled at its default so queued jobs survive a cache restart.

Find the private network the template created so you can attach the worker to it:

bash
openstack network list

Note the network name whose subnet matches the cache CIDR (10.20.0.0/24 by default). The worker joins this network in step 3.

Step 2: Create the result bucket and S3 credentials#

The worker writes each result to Object Storage. Create a bucket and a set of S3 credentials it uses.

Create the bucket by following Create a container, and name it gpu-results. Then generate an access key and secret by following Create S3 credentials. Record the access key, the secret, and the S3 endpoint URL for your region.

You pass these to the worker as environment variables in step 4. They stay on the worker instance you own and never travel to the GPU provider.

Step 3: Provision the worker instance#

Create a CPU instance on the cache's private network so it reaches the queue over private addressing. Give it a floating IP for SSH; outbound calls to the GPU API and Object Storage route through the network's router.

Create a security group that allows SSH, then launch the worker:

bash
openstack security group create worker-ssh --description "SSH to the GPU control-plane worker"
openstack security group rule create --proto tcp --dst-port 22 --remote-ip 0.0.0.0/0 worker-ssh

openstack server create \
  --flavor s1a.small \
  --image Ubuntu-24.04 \
  --key-name YOUR_KEY_NAME \
  --network CACHE_PRIVATE_NETWORK \
  --security-group worker-ssh \
  gpu-control-plane-worker

Allocate a floating IP and attach it so you can connect:

bash
WORKER_FIP=$(openstack floating ip create PublicStatic -f value -c floating_ip_address)
openstack server add floating ip gpu-control-plane-worker "$WORKER_FIP"
echo "$WORKER_FIP"

The worker now sits on the same private network as the cache and has outbound access to the GPU provider and to Object Storage.

Step 4: Install dependencies and set credentials#

Connect to the worker with the private key that matches key_name:

bash
ssh -i ~/.ssh/YOUR_KEY -o StrictHostKeyChecking=accept-new ubuntu@"$WORKER_FIP"

Install Python and the three libraries the worker uses:

bash
sudo apt-get update && sudo apt-get install -y python3-venv
python3 -m venv ~/worker-env
~/worker-env/bin/pip install redis requests boto3

Write the credentials and configuration to an environment file the worker reads. Keep this file on the instance; the GPU key and the S3 secret live here, on the control plane you own, not in the queue and not in client code:

bash
cat > ~/worker.env <<'EOF'
REDIS_URL=redis://:REDIS_PASSWORD@CACHE_PRIVATE_IP:6379/0
QUEUE_KEY=gpu:jobs
GPU_API_BASE=https://YOUR_GPU_PROVIDER/v1
GPU_API_KEY=YOUR_GPU_PROVIDER_KEY
S3_ENDPOINT=https://YOUR_OBJECT_STORAGE_ENDPOINT
AWS_ACCESS_KEY_ID=YOUR_S3_ACCESS_KEY
AWS_SECRET_ACCESS_KEY=YOUR_S3_SECRET_KEY
RESULT_BUCKET=gpu-results
MAX_ATTEMPTS=3
DAILY_JOB_CAP=1000
EOF
chmod 600 ~/worker.env

Replace the placeholders with the cache password and private IP from step 1, the GPU provider base URL and key from your provider, and the S3 endpoint, access key, and secret from step 2. DAILY_JOB_CAP is a per-day job count the worker enforces before it submits, so a runaway producer cannot drive unbounded spend on the GPU provider. Restrict it with file permissions so other accounts on the instance cannot read it.

For a stricter secret boundary, move these values out of a plain file with How to store application secrets and inject them at runtime.

Step 5: Write the worker#

The worker pops a job, submits it to the GPU API, polls until the result is ready, and stores the output. Create worker.py:

Python
import json
import os
import time

import boto3
import redis
import requests

REDIS_URL = os.environ["REDIS_URL"]
QUEUE_KEY = os.environ.get("QUEUE_KEY", "gpu:jobs")
GPU_API_BASE = os.environ["GPU_API_BASE"].rstrip("/")
GPU_API_KEY = os.environ["GPU_API_KEY"]
RESULT_BUCKET = os.environ["RESULT_BUCKET"]
MAX_ATTEMPTS = int(os.environ.get("MAX_ATTEMPTS", "3"))
DAILY_JOB_CAP = int(os.environ.get("DAILY_JOB_CAP", "1000"))
POLL_INTERVAL = float(os.environ.get("POLL_INTERVAL", "5"))
POLL_TIMEOUT = float(os.environ.get("POLL_TIMEOUT", "600"))

queue = redis.Redis.from_url(REDIS_URL)
storage = boto3.client("s3", endpoint_url=os.environ["S3_ENDPOINT"])
http = requests.Session()
http.headers["Authorization"] = f"Bearer {GPU_API_KEY}"


def within_budget():
    day = time.strftime("%Y%m%d")
    used = queue.incr(f"gpu:spend:{day}")
    queue.expire(f"gpu:spend:{day}", 172800)
    return used <= DAILY_JOB_CAP


def submit(job):
    response = http.post(f"{GPU_API_BASE}/jobs", json=job["input"], timeout=30)
    response.raise_for_status()
    return response.json()["id"]


def poll(remote_id):
    deadline = time.monotonic() + POLL_TIMEOUT
    while time.monotonic() < deadline:
        response = http.get(f"{GPU_API_BASE}/jobs/{remote_id}", timeout=30)
        response.raise_for_status()
        body = response.json()
        if body["status"] == "succeeded":
            return body["output"]
        if body["status"] == "failed":
            raise RuntimeError(f"remote job {remote_id} failed")
        time.sleep(POLL_INTERVAL)
    raise TimeoutError(f"remote job {remote_id} did not finish in time")


def store(job, output):
    storage.put_object(
        Bucket=RESULT_BUCKET,
        Key=job["output_key"],
        Body=json.dumps(output).encode(),
        ContentType="application/json",
    )
    return job["output_key"]


def process(job):
    remote_id = submit(job)
    output = poll(remote_id)
    return store(job, output)


def main():
    print("worker ready, waiting for jobs")
    while True:
        item = queue.blpop(QUEUE_KEY, timeout=5)
        if item is None:
            continue
        job = json.loads(item[1])
        attempt = job.get("attempt", 1)
        if not within_budget():
            queue.rpush(f"{QUEUE_KEY}:held", json.dumps(job))
            print(f"job {job['id']} held: daily cap reached")
            time.sleep(60)
            continue
        try:
            key = process(job)
            print(f"job {job['id']} done, result at {key}")
        except Exception as error:
            if attempt < MAX_ATTEMPTS:
                job["attempt"] = attempt + 1
                backoff = 2 ** attempt
                print(f"job {job['id']} failed ({error}); retry in {backoff}s")
                time.sleep(backoff)
                queue.rpush(QUEUE_KEY, json.dumps(job))
            else:
                queue.rpush(
                    f"{QUEUE_KEY}:dead",
                    json.dumps({"job": job, "error": str(error)}),
                )
                print(f"job {job['id']} exhausted retries, moved to dead-letter")


if __name__ == "__main__":
    main()

The worker handles the three failure cases a remote backend creates: a submit or poll error retries with exponential backoff and re-enqueues the job, a job that exhausts MAX_ATTEMPTS moves to a gpu:jobs:dead dead-letter list you can inspect, and a job that would exceed the daily cap moves to a gpu:jobs:held list instead of submitting. The credentials and the cap stay in the worker's environment.

The submit and poll functions match a provider that returns a job ID and a status you poll. For a synchronous API, such as an OpenAI-compatible chat-completions endpoint, replace process with a single POST and store the response directly; the queue, retry, budget, and storage logic stay the same.

Step 6: Enqueue jobs and run the worker#

Write a small producer that pushes jobs onto the queue. Create enqueue.py:

Python
import json
import os
import sys
import uuid

import redis

queue = redis.Redis.from_url(os.environ["REDIS_URL"])
QUEUE_KEY = os.environ.get("QUEUE_KEY", "gpu:jobs")


def enqueue(prompt):
    job_id = uuid.uuid4().hex
    job = {
        "id": job_id,
        "input": {"prompt": prompt},
        "output_key": f"results/{job_id}.json",
        "attempt": 1,
    }
    queue.rpush(QUEUE_KEY, json.dumps(job))
    print(job_id)


if __name__ == "__main__":
    for prompt in sys.argv[1:]:
        enqueue(prompt)

Load the environment and start the worker in one terminal:

bash
set -a && source ~/worker.env && set +a
~/worker-env/bin/python worker.py

The worker prints worker ready, waiting for jobs. In a second SSH session, load the environment and enqueue two jobs:

bash
set -a && source ~/worker.env && set +a
~/worker-env/bin/python enqueue.py "a short poem about queues" "a haiku about caching"

Each call prints a job ID. The worker terminal then logs job <id> done, result at results/<id>.json for each one as it submits the job, polls the provider, and stores the output.

Step 7: Fan out across several workers#

One worker processes jobs one at a time. Because each worker uses a blocking BLPOP, several workers draw from the same queue without coordination, and each job goes to exactly one worker. Start three workers on the instance for the demo:

bash
for i in 1 2 3; do ~/worker-env/bin/python worker.py & done

Enqueue a batch and watch the three workers share it. Press Ctrl+C and run kill %1 %2 %3 to stop the background workers when you finish.

For a durable setup, run each worker under a systemd template unit so the host restarts them on failure and on reboot:

ini
# /etc/systemd/system/[email protected]
[Unit]
Description=GPU control-plane worker %i
After=network-online.target

[Service]
User=ubuntu
EnvironmentFile=/home/ubuntu/worker.env
ExecStart=/home/ubuntu/worker-env/bin/python /home/ubuntu/worker.py
Restart=always

[Install]
WantedBy=multi-user.target

Enable three instances of the unit:

bash
sudo systemctl daemon-reload
sudo systemctl enable --now gpu-worker@1 gpu-worker@2 gpu-worker@3

When one instance saturates, add more worker instances on the same private network with the step 3 commands. Every worker reads the same queue, so throughput scales with the worker count up to the rate limit of the GPU provider.

Step 8: Confirm results in Object Storage#

List the results bucket from your workstation to confirm the worker wrote each output. With the S3 credentials configured for the AWS CLI:

bash
aws --endpoint-url https://YOUR_OBJECT_STORAGE_ENDPOINT \
  s3 ls s3://gpu-results/results/

You see one object per completed job, keyed by job ID. Download one to inspect it:

bash
aws --endpoint-url https://YOUR_OBJECT_STORAGE_ENDPOINT \
  s3 cp s3://gpu-results/results/JOB_ID.json -

The object holds the provider output the worker stored. Jobs that exhausted their retries are in the gpu:jobs:dead list on the cache; read them with redis-cli LRANGE gpu:jobs:dead 0 -1 from a host on the private network.

Step 9: Tear down#

Remove the worker, then the queue.

Delete the worker instance, its floating IP, and the security group:

bash
openstack server delete gpu-control-plane-worker
openstack floating ip delete "$WORKER_FIP"
openstack security group delete worker-ssh

Destroy the queue from the directory where you applied the redis-cache template:

bash
tofu destroy

Empty and delete the gpu-results bucket from the Console or the CLI when you no longer need the stored results.

What you built#

  • Stood up a queue with the redis-cache template on a private network
  • Created an Object Storage bucket and S3 credentials for the results
  • Provisioned a CPU worker on the queue's private network with outbound access to a GPU provider
  • Wrote a worker that pops a job, submits it to a third-party GPU or model API, polls for completion, and stores the result in Object Storage
  • Added retry with backoff, a dead-letter queue, and a daily job cap, with the GPU credentials and the cap on the control plane you own
  • Fanned out across several workers drawing from the same queue

The queue and the workers are the always-on control plane on Quake AI CPU compute. The GPU compute runs on the provider you point the worker at, and you repoint it by changing GPU_API_BASE and GPU_API_KEY in one environment file.

Next steps#

Clean up#

If any resources from this walkthrough remain, delete the worker instance and its floating IP with the OpenStack CLI and run tofu destroy against the queue so the instances, volumes, network, and floating IP stop holding quota.

Quick answers

Was this page helpful?