# Deploy a bursty-GPU control plane with a queue and workers

Source: https://docs.quake.ai/resources/deployments/deploy-gpu-control-plane
Markdown: https://docs.quake.ai/resources/deployments/deploy-gpu-control-plane.md
> Run the durable control plane on Quake AI CPU compute: a Redis or Valkey queue and a worker process accept jobs, submit each one to a third-party GPU or model API, poll for completion, retry on failure, and land results in Object Storage. The queue and workers stay always-on and cheap; the GPU work runs on a provider you choose.

---

# Deploy a bursty-GPU control plane with a queue and workers

Stand up a bursty-GPU control plane on Quake AI CPU compute using the [validated OpenTofu template](/docs/platform/validation#how-infrastructure-templates-are-checked) `redis-cache` for the job queue. A Redis or Valkey queue holds jobs, a worker process on a CPU instance pops each job, submits it to a third-party GPU or model API, polls until the result is ready, and writes the output to Object Storage. The queue and the worker run on instances you own and keep running cheaply; the heavy GPU compute runs on whichever provider you point the worker at.

This is the brain-on-CPU, muscle-elsewhere shape from [AI inference and RAG pipelines](/resources/solutions/ai-inference-rag), made concrete as a job pipeline. The durable parts, the queue, the worker loop, the retry policy, the credentials, and the result store, live on the control plane you operate. The GPU backend is a callout the worker makes, so you can repoint it at a different provider without changing the pipeline around it. [OpenClaw](/docs/compute/apps/openclaw) runs the same always-on-agent-on-CPU shape as a one-click Console App.

<Figure size="md" caption="What you will build: an enqueue step pushes jobs to a Redis or Valkey queue on Quake AI; a CPU worker pops each job, submits it to a GPU provider elsewhere, polls for the result, and writes it to Object Storage">

```d2
direction: right

producer: Your app or workstation {
  enqueue: Enqueue job
}

cloud: Quake AI control plane {
  queue: Redis / Valkey queue\n(private IP)
  worker: Worker\n(CPU instance)
  bucket: Object Storage\nresults {shape: cylinder}
  worker -> queue: BLPOP job
  worker -> bucket: store result
}

gpu: GPU provider (elsewhere) {
  api: Inference or GPU API
}

producer.enqueue -> cloud.queue: RPUSH
cloud.worker -> gpu.api: submit and poll
```

</Figure>

<PricingCompanion
  components={[
    { kind: "template", slug: "redis-cache", required: true },
  ]}
/>

The queue is the `redis-cache` template above. The worker instance, the Object Storage bucket, and the worker code are the custom layer you add on top in the steps below.

## Prerequisites

Before you start, confirm you have:

- A Quake AI account with [application credentials](/docs/tools/generate-app-credentials) and the [OpenStack CLI](/docs/tools/openstack-cli) configured
- OpenTofu installed locally ([installation guide](https://opentofu.org/docs/intro/install/)) for the queue step
- An SSH key pair that already exists in your project; record its name for the `key_name` variables
- An API key for one third-party GPU or model provider with an HTTP API, and the base URL of that API
- Enough project quota for two `s1a` instances, two volumes, one private network, a router, two security groups, and one floating IP
- Basic familiarity with [Object Storage on Quake AI](/docs/object)

This pipeline runs CPU-only on Quake AI. The GPU compute runs on the external provider you choose; Quake AI flavors are AMD EPYC with no GPU option ([compute FAQ](/docs/compute/faq)). The pattern fits batch inference, image and audio generation, and any job you submit to a remote accelerator and collect later.

## Step 1: Stand up the queue

The queue is a Redis or Valkey instance on a private network. Deploy it with the `redis-cache` template by following [Deploy a Redis or Valkey cache with the redis-cache template](/resources/deployments/deploy-redis-cache-template), then return here. That walkthrough provisions the private network, the security group, the optional persistence volume, and the cache instance, and it leaves the cache reachable only from the private subnet.

When the cache stack is up, capture three values from its `tofu output`:

```bash
CACHE_PRIVATE_IP=$(tofu output -raw private_ip)
echo "$CACHE_PRIVATE_IP"
```

Record the cache's private IP, the private network name, and the password you set. The worker reaches the cache at `redis://:REDIS_PASSWORD@CACHE_PRIVATE_IP:6379/0` over the private network. Leave `persistence_enabled` at its default so queued jobs survive a cache restart.

Find the private network the template created so you can attach the worker to it:

```bash
openstack network list
```

Note the network name whose subnet matches the cache CIDR (`10.20.0.0/24` by default). The worker joins this network in step 3.

## Step 2: Create the result bucket and S3 credentials

The worker writes each result to Object Storage. Create a bucket and a set of S3 credentials it uses.

Create the bucket by following [Create a container](/docs/object/how-to/create-container), and name it `gpu-results`. Then generate an access key and secret by following [Create S3 credentials](/docs/object/how-to/create-s3-credentials). Record the access key, the secret, and the S3 endpoint URL for your region.

You pass these to the worker as environment variables in step 4. They stay on the worker instance you own and never travel to the GPU provider.

## Step 3: Provision the worker instance

Create a CPU instance on the cache's private network so it reaches the queue over private addressing. Give it a floating IP for SSH; outbound calls to the GPU API and Object Storage route through the network's router.

Create a security group that allows SSH, then launch the worker:

```bash
openstack security group create worker-ssh --description "SSH to the GPU control-plane worker"
openstack security group rule create --proto tcp --dst-port 22 --remote-ip 0.0.0.0/0 worker-ssh

openstack server create \
  --flavor s1a.small \
  --image Ubuntu-24.04 \
  --key-name YOUR_KEY_NAME \
  --network CACHE_PRIVATE_NETWORK \
  --security-group worker-ssh \
  gpu-control-plane-worker
```

Allocate a floating IP and attach it so you can connect:

```bash
WORKER_FIP=$(openstack floating ip create PublicStatic -f value -c floating_ip_address)
openstack server add floating ip gpu-control-plane-worker "$WORKER_FIP"
echo "$WORKER_FIP"
```

The worker now sits on the same private network as the cache and has outbound access to the GPU provider and to Object Storage.

## Step 4: Install dependencies and set credentials

Connect to the worker with the private key that matches `key_name`:

```bash
ssh -i ~/.ssh/YOUR_KEY -o StrictHostKeyChecking=accept-new ubuntu@"$WORKER_FIP"
```

Install Python and the three libraries the worker uses:

```bash
sudo apt-get update && sudo apt-get install -y python3-venv
python3 -m venv ~/worker-env
~/worker-env/bin/pip install redis requests boto3
```

Write the credentials and configuration to an environment file the worker reads. Keep this file on the instance; the GPU key and the S3 secret live here, on the control plane you own, not in the queue and not in client code:

```bash
cat > ~/worker.env <<'EOF'
REDIS_URL=redis://:REDIS_PASSWORD@CACHE_PRIVATE_IP:6379/0
QUEUE_KEY=gpu:jobs
GPU_API_BASE=https://YOUR_GPU_PROVIDER/v1
GPU_API_KEY=YOUR_GPU_PROVIDER_KEY
S3_ENDPOINT=https://YOUR_OBJECT_STORAGE_ENDPOINT
AWS_ACCESS_KEY_ID=YOUR_S3_ACCESS_KEY
AWS_SECRET_ACCESS_KEY=YOUR_S3_SECRET_KEY
RESULT_BUCKET=gpu-results
MAX_ATTEMPTS=3
DAILY_JOB_CAP=1000
EOF
chmod 600 ~/worker.env
```

Replace the placeholders with the cache password and private IP from step 1, the GPU provider base URL and key from your provider, and the S3 endpoint, access key, and secret from step 2. `DAILY_JOB_CAP` is a per-day job count the worker enforces before it submits, so a runaway producer cannot drive unbounded spend on the GPU provider. Restrict it with file permissions so other accounts on the instance cannot read it.

For a stricter secret boundary, move these values out of a plain file with [How to store application secrets and inject them at runtime](/docs/security/how-to/inject-app-secrets).

## Step 5: Write the worker

The worker pops a job, submits it to the GPU API, polls until the result is ready, and stores the output. Create `worker.py`:

```python
import json
import os
import time

import boto3
import redis
import requests

REDIS_URL = os.environ["REDIS_URL"]
QUEUE_KEY = os.environ.get("QUEUE_KEY", "gpu:jobs")
GPU_API_BASE = os.environ["GPU_API_BASE"].rstrip("/")
GPU_API_KEY = os.environ["GPU_API_KEY"]
RESULT_BUCKET = os.environ["RESULT_BUCKET"]
MAX_ATTEMPTS = int(os.environ.get("MAX_ATTEMPTS", "3"))
DAILY_JOB_CAP = int(os.environ.get("DAILY_JOB_CAP", "1000"))
POLL_INTERVAL = float(os.environ.get("POLL_INTERVAL", "5"))
POLL_TIMEOUT = float(os.environ.get("POLL_TIMEOUT", "600"))

queue = redis.Redis.from_url(REDIS_URL)
storage = boto3.client("s3", endpoint_url=os.environ["S3_ENDPOINT"])
http = requests.Session()
http.headers["Authorization"] = f"Bearer {GPU_API_KEY}"


def within_budget():
    day = time.strftime("%Y%m%d")
    used = queue.incr(f"gpu:spend:{day}")
    queue.expire(f"gpu:spend:{day}", 172800)
    return used <= DAILY_JOB_CAP


def submit(job):
    response = http.post(f"{GPU_API_BASE}/jobs", json=job["input"], timeout=30)
    response.raise_for_status()
    return response.json()["id"]


def poll(remote_id):
    deadline = time.monotonic() + POLL_TIMEOUT
    while time.monotonic() < deadline:
        response = http.get(f"{GPU_API_BASE}/jobs/{remote_id}", timeout=30)
        response.raise_for_status()
        body = response.json()
        if body["status"] == "succeeded":
            return body["output"]
        if body["status"] == "failed":
            raise RuntimeError(f"remote job {remote_id} failed")
        time.sleep(POLL_INTERVAL)
    raise TimeoutError(f"remote job {remote_id} did not finish in time")


def store(job, output):
    storage.put_object(
        Bucket=RESULT_BUCKET,
        Key=job["output_key"],
        Body=json.dumps(output).encode(),
        ContentType="application/json",
    )
    return job["output_key"]


def process(job):
    remote_id = submit(job)
    output = poll(remote_id)
    return store(job, output)


def main():
    print("worker ready, waiting for jobs")
    while True:
        item = queue.blpop(QUEUE_KEY, timeout=5)
        if item is None:
            continue
        job = json.loads(item[1])
        attempt = job.get("attempt", 1)
        if not within_budget():
            queue.rpush(f"{QUEUE_KEY}:held", json.dumps(job))
            print(f"job {job['id']} held: daily cap reached")
            time.sleep(60)
            continue
        try:
            key = process(job)
            print(f"job {job['id']} done, result at {key}")
        except Exception as error:
            if attempt < MAX_ATTEMPTS:
                job["attempt"] = attempt + 1
                backoff = 2 ** attempt
                print(f"job {job['id']} failed ({error}); retry in {backoff}s")
                time.sleep(backoff)
                queue.rpush(QUEUE_KEY, json.dumps(job))
            else:
                queue.rpush(
                    f"{QUEUE_KEY}:dead",
                    json.dumps({"job": job, "error": str(error)}),
                )
                print(f"job {job['id']} exhausted retries, moved to dead-letter")


if __name__ == "__main__":
    main()
```

The worker handles the three failure cases a remote backend creates: a submit or poll error retries with exponential backoff and re-enqueues the job, a job that exhausts `MAX_ATTEMPTS` moves to a `gpu:jobs:dead` dead-letter list you can inspect, and a job that would exceed the daily cap moves to a `gpu:jobs:held` list instead of submitting. The credentials and the cap stay in the worker's environment.

The `submit` and `poll` functions match a provider that returns a job ID and a status you poll. For a synchronous API, such as an OpenAI-compatible chat-completions endpoint, replace `process` with a single `POST` and store the response directly; the queue, retry, budget, and storage logic stay the same.

## Step 6: Enqueue jobs and run the worker

Write a small producer that pushes jobs onto the queue. Create `enqueue.py`:

```python
import json
import os
import sys
import uuid

import redis

queue = redis.Redis.from_url(os.environ["REDIS_URL"])
QUEUE_KEY = os.environ.get("QUEUE_KEY", "gpu:jobs")


def enqueue(prompt):
    job_id = uuid.uuid4().hex
    job = {
        "id": job_id,
        "input": {"prompt": prompt},
        "output_key": f"results/{job_id}.json",
        "attempt": 1,
    }
    queue.rpush(QUEUE_KEY, json.dumps(job))
    print(job_id)


if __name__ == "__main__":
    for prompt in sys.argv[1:]:
        enqueue(prompt)
```

Load the environment and start the worker in one terminal:

```bash
set -a && source ~/worker.env && set +a
~/worker-env/bin/python worker.py
```

The worker prints `worker ready, waiting for jobs`. In a second SSH session, load the environment and enqueue two jobs:

```bash
set -a && source ~/worker.env && set +a
~/worker-env/bin/python enqueue.py "a short poem about queues" "a haiku about caching"
```

Each call prints a job ID. The worker terminal then logs `job <id> done, result at results/<id>.json` for each one as it submits the job, polls the provider, and stores the output.

## Step 7: Fan out across several workers

One worker processes jobs one at a time. Because each worker uses a blocking `BLPOP`, several workers draw from the same queue without coordination, and each job goes to exactly one worker. Start three workers on the instance for the demo:

```bash
for i in 1 2 3; do ~/worker-env/bin/python worker.py & done
```

Enqueue a batch and watch the three workers share it. Press `Ctrl+C` and run `kill %1 %2 %3` to stop the background workers when you finish.

For a durable setup, run each worker under a systemd template unit so the host restarts them on failure and on reboot:

```ini
# /etc/systemd/system/gpu-worker@.service
[Unit]
Description=GPU control-plane worker %i
After=network-online.target

[Service]
User=ubuntu
EnvironmentFile=/home/ubuntu/worker.env
ExecStart=/home/ubuntu/worker-env/bin/python /home/ubuntu/worker.py
Restart=always

[Install]
WantedBy=multi-user.target
```

Enable three instances of the unit:

```bash
sudo systemctl daemon-reload
sudo systemctl enable --now gpu-worker@1 gpu-worker@2 gpu-worker@3
```

When one instance saturates, add more worker instances on the same private network with the step 3 commands. Every worker reads the same queue, so throughput scales with the worker count up to the rate limit of the GPU provider.

## Step 8: Confirm results in Object Storage

List the results bucket from your workstation to confirm the worker wrote each output. With the S3 credentials configured for the AWS CLI:

```bash
aws --endpoint-url https://YOUR_OBJECT_STORAGE_ENDPOINT \
  s3 ls s3://gpu-results/results/
```

You see one object per completed job, keyed by job ID. Download one to inspect it:

```bash
aws --endpoint-url https://YOUR_OBJECT_STORAGE_ENDPOINT \
  s3 cp s3://gpu-results/results/JOB_ID.json -
```

The object holds the provider output the worker stored. Jobs that exhausted their retries are in the `gpu:jobs:dead` list on the cache; read them with `redis-cli LRANGE gpu:jobs:dead 0 -1` from a host on the private network.

## Step 9: Tear down

Remove the worker, then the queue.

Delete the worker instance, its floating IP, and the security group:

```bash
openstack server delete gpu-control-plane-worker
openstack floating ip delete "$WORKER_FIP"
openstack security group delete worker-ssh
```

Destroy the queue from the directory where you applied the `redis-cache` template:

```bash
tofu destroy
```

Empty and delete the `gpu-results` bucket from the Console or the CLI when you no longer need the stored results.

## What you built

- **Stood up a queue** with the `redis-cache` template on a private network
- **Created an Object Storage bucket and S3 credentials** for the results
- **Provisioned a CPU worker** on the queue's private network with outbound access to a GPU provider
- **Wrote a worker** that pops a job, submits it to a third-party GPU or model API, polls for completion, and stores the result in Object Storage
- **Added retry with backoff, a dead-letter queue, and a daily job cap**, with the GPU credentials and the cap on the control plane you own
- **Fanned out across several workers** drawing from the same queue

The queue and the workers are the always-on control plane on Quake AI CPU compute. The GPU compute runs on the provider you point the worker at, and you repoint it by changing `GPU_API_BASE` and `GPU_API_KEY` in one environment file.

## Next steps

- [AI inference and RAG pipelines](/resources/solutions/ai-inference-rag): the solution pattern this pipeline fits into
- [Deploy an inference gateway with OpenTofu](/resources/deployments/deploy-inference-gateway-template): put one stable endpoint in front of several model backends, then point the worker's `GPU_API_BASE` at it
- [Redis / Valkey cache template](/resources/iac-templates/redis-cache): the queue's parameters and resource map
- [Upload objects to Object Storage](/docs/object/how-to/upload-file): the storage tier the worker writes to
- [How to store application secrets and inject them at runtime](/docs/security/how-to/inject-app-secrets): move the worker credentials out of a plain environment file

## Clean up

If any resources from this walkthrough remain, delete the worker instance and its floating IP with the OpenStack CLI and run `tofu destroy` against the queue so the instances, volumes, network, and floating IP stop holding quota.
