Deploy a bursty-GPU control plane with a queue and workers
Deployment
Coming from another cloud?
▸AWS·SQS
This Quake AI feature maps to AWS’s SQS.
▸Google Cloud·Pubsub
This Quake AI feature maps to Google Cloud’s Pubsub.
Deploy a bursty-GPU control plane with a queue and workers
Stand up a bursty-GPU control plane on Quake AI CPU compute using the validated OpenTofu templateredis-cache for the job queue. A Redis or Valkey queue holds jobs, a worker process on a CPU instance pops each job, submits it to a third-party GPU or model API, polls until the result is ready, and writes the output to Object Storage. The queue and the worker run on instances you own and keep running cheaply; the heavy GPU compute runs on whichever provider you point the worker at.
This is the brain-on-CPU, muscle-elsewhere shape from AI inference and RAG pipelines, made concrete as a job pipeline. The durable parts, the queue, the worker loop, the retry policy, the credentials, and the result store, live on the control plane you operate. The GPU backend is a callout the worker makes, so you can repoint it at a different provider without changing the pipeline around it. OpenClaw runs the same always-on-agent-on-CPU shape as a one-click Console App.
Click to zoom
What you will build: an enqueue step pushes jobs to a Redis or Valkey queue on Quake AI; a CPU worker pops each job, submits it to a GPU provider elsewhere, polls for the result, and writes it to Object Storage
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Cache
s1a.small · 2 shared vCPU, 2 GiB RAM, 0.5 Gbps
$16.50/mo
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
$0.00
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
$0.00
Pricing data last validated: . For current rates, check quake.ai/pricing.
The queue is the redis-cache template above. The worker instance, the Object Storage bucket, and the worker code are the custom layer you add on top in the steps below.
This pipeline runs CPU-only on Quake AI. The GPU compute runs on the external provider you choose; Quake AI flavors are AMD EPYC with no GPU option (compute FAQ). The pattern fits batch inference, image and audio generation, and any job you submit to a remote accelerator and collect later.
The queue is a Redis or Valkey instance on a private network. Deploy it with the redis-cache template by following Deploy a Redis or Valkey cache with the redis-cache template, then return here. That walkthrough provisions the private network, the security group, the optional persistence volume, and the cache instance, and it leaves the cache reachable only from the private subnet.
When the cache stack is up, capture three values from its tofu output:
Record the cache's private IP, the private network name, and the password you set. The worker reaches the cache at redis://:REDIS_PASSWORD@CACHE_PRIVATE_IP:6379/0 over the private network. Leave persistence_enabled at its default so queued jobs survive a cache restart.
Find the private network the template created so you can attach the worker to it:
bash
openstack network list
Note the network name whose subnet matches the cache CIDR (10.20.0.0/24 by default). The worker joins this network in step 3.
Step 2: Create the result bucket and S3 credentials#
The worker writes each result to Object Storage. Create a bucket and a set of S3 credentials it uses.
Create the bucket by following Create a container, and name it gpu-results. Then generate an access key and secret by following Create S3 credentials. Record the access key, the secret, and the S3 endpoint URL for your region.
You pass these to the worker as environment variables in step 4. They stay on the worker instance you own and never travel to the GPU provider.
Create a CPU instance on the cache's private network so it reaches the queue over private addressing. Give it a floating IP for SSH; outbound calls to the GPU API and Object Storage route through the network's router.
Create a security group that allows SSH, then launch the worker:
bash
openstack security group create worker-ssh --description "SSH to the GPU control-plane worker"openstack security group rule create --proto tcp --dst-port 22 --remote-ip 0.0.0.0/0 worker-sshopenstack server create \ --flavor s1a.small \ --image Ubuntu-24.04 \ --key-name YOUR_KEY_NAME \ --network CACHE_PRIVATE_NETWORK \ --security-group worker-ssh \ gpu-control-plane-worker
Allocate a floating IP and attach it so you can connect:
bash
WORKER_FIP=$(openstack floating ip create PublicStatic -f value -c floating_ip_address)openstack server add floating ip gpu-control-plane-worker "$WORKER_FIP"echo "$WORKER_FIP"
The worker now sits on the same private network as the cache and has outbound access to the GPU provider and to Object Storage.
Write the credentials and configuration to an environment file the worker reads. Keep this file on the instance; the GPU key and the S3 secret live here, on the control plane you own, not in the queue and not in client code:
Replace the placeholders with the cache password and private IP from step 1, the GPU provider base URL and key from your provider, and the S3 endpoint, access key, and secret from step 2. DAILY_JOB_CAP is a per-day job count the worker enforces before it submits, so a runaway producer cannot drive unbounded spend on the GPU provider. Restrict it with file permissions so other accounts on the instance cannot read it.
The worker pops a job, submits it to the GPU API, polls until the result is ready, and stores the output. Create worker.py:
Python
import jsonimport osimport timeimport boto3import redisimport requestsREDIS_URL = os.environ["REDIS_URL"]QUEUE_KEY = os.environ.get("QUEUE_KEY", "gpu:jobs")GPU_API_BASE = os.environ["GPU_API_BASE"].rstrip("/")GPU_API_KEY = os.environ["GPU_API_KEY"]RESULT_BUCKET = os.environ["RESULT_BUCKET"]MAX_ATTEMPTS = int(os.environ.get("MAX_ATTEMPTS", "3"))DAILY_JOB_CAP = int(os.environ.get("DAILY_JOB_CAP", "1000"))POLL_INTERVAL = float(os.environ.get("POLL_INTERVAL", "5"))POLL_TIMEOUT = float(os.environ.get("POLL_TIMEOUT", "600"))queue = redis.Redis.from_url(REDIS_URL)storage = boto3.client("s3", endpoint_url=os.environ["S3_ENDPOINT"])http = requests.Session()http.headers["Authorization"] = f"Bearer {GPU_API_KEY}"def within_budget(): day = time.strftime("%Y%m%d") used = queue.incr(f"gpu:spend:{day}") queue.expire(f"gpu:spend:{day}", 172800) return used <= DAILY_JOB_CAPdef submit(job): response = http.post(f"{GPU_API_BASE}/jobs", json=job["input"], timeout=30) response.raise_for_status() return response.json()["id"]def poll(remote_id): deadline = time.monotonic() + POLL_TIMEOUT while time.monotonic() < deadline: response = http.get(f"{GPU_API_BASE}/jobs/{remote_id}", timeout=30) response.raise_for_status() body = response.json() if body["status"] == "succeeded": return body["output"] if body["status"] == "failed": raise RuntimeError(f"remote job {remote_id} failed") time.sleep(POLL_INTERVAL) raise TimeoutError(f"remote job {remote_id} did not finish in time")def store(job, output): storage.put_object( Bucket=RESULT_BUCKET, Key=job["output_key"], Body=json.dumps(output).encode(), ContentType="application/json", ) return job["output_key"]def process(job): remote_id = submit(job) output = poll(remote_id) return store(job, output)def main(): print("worker ready, waiting for jobs") while True: item = queue.blpop(QUEUE_KEY, timeout=5) if item is None: continue job = json.loads(item[1]) attempt = job.get("attempt", 1) if not within_budget(): queue.rpush(f"{QUEUE_KEY}:held", json.dumps(job)) print(f"job {job['id']} held: daily cap reached") time.sleep(60) continue try: key = process(job) print(f"job {job['id']} done, result at {key}") except Exception as error: if attempt < MAX_ATTEMPTS: job["attempt"] = attempt + 1 backoff = 2 ** attempt print(f"job {job['id']} failed ({error}); retry in {backoff}s") time.sleep(backoff) queue.rpush(QUEUE_KEY, json.dumps(job)) else: queue.rpush( f"{QUEUE_KEY}:dead", json.dumps({"job": job, "error": str(error)}), ) print(f"job {job['id']} exhausted retries, moved to dead-letter")if __name__ == "__main__": main()
The worker handles the three failure cases a remote backend creates: a submit or poll error retries with exponential backoff and re-enqueues the job, a job that exhausts MAX_ATTEMPTS moves to a gpu:jobs:dead dead-letter list you can inspect, and a job that would exceed the daily cap moves to a gpu:jobs:held list instead of submitting. The credentials and the cap stay in the worker's environment.
The submit and poll functions match a provider that returns a job ID and a status you poll. For a synchronous API, such as an OpenAI-compatible chat-completions endpoint, replace process with a single POST and store the response directly; the queue, retry, budget, and storage logic stay the same.
Load the environment and start the worker in one terminal:
bash
set -a && source ~/worker.env && set +a~/worker-env/bin/python worker.py
The worker prints worker ready, waiting for jobs. In a second SSH session, load the environment and enqueue two jobs:
bash
set -a && source ~/worker.env && set +a~/worker-env/bin/python enqueue.py "a short poem about queues" "a haiku about caching"
Each call prints a job ID. The worker terminal then logs job <id> done, result at results/<id>.json for each one as it submits the job, polls the provider, and stores the output.
One worker processes jobs one at a time. Because each worker uses a blocking BLPOP, several workers draw from the same queue without coordination, and each job goes to exactly one worker. Start three workers on the instance for the demo:
bash
for i in 1 2 3; do ~/worker-env/bin/python worker.py & done
Enqueue a batch and watch the three workers share it. Press Ctrl+C and run kill %1 %2 %3 to stop the background workers when you finish.
For a durable setup, run each worker under a systemd template unit so the host restarts them on failure and on reboot:
When one instance saturates, add more worker instances on the same private network with the step 3 commands. Every worker reads the same queue, so throughput scales with the worker count up to the rate limit of the GPU provider.
The object holds the provider output the worker stored. Jobs that exhausted their retries are in the gpu:jobs:dead list on the cache; read them with redis-cli LRANGE gpu:jobs:dead 0 -1 from a host on the private network.
Stood up a queue with the redis-cache template on a private network
Created an Object Storage bucket and S3 credentials for the results
Provisioned a CPU worker on the queue's private network with outbound access to a GPU provider
Wrote a worker that pops a job, submits it to a third-party GPU or model API, polls for completion, and stores the result in Object Storage
Added retry with backoff, a dead-letter queue, and a daily job cap, with the GPU credentials and the cap on the control plane you own
Fanned out across several workers drawing from the same queue
The queue and the workers are the always-on control plane on Quake AI CPU compute. The GPU compute runs on the provider you point the worker at, and you repoint it by changing GPU_API_BASE and GPU_API_KEY in one environment file.
If any resources from this walkthrough remain, delete the worker instance and its floating IP with the OpenStack CLI and run tofu destroy against the queue so the instances, volumes, network, and floating IP stop holding quota.