Creator-AI inference worker
Creator-AI inference worker
This pattern composes Compute and Object Storage.
What this template deploys#
- Private network and router for the inference instance
- One Object Storage bucket for input media, generated captions, and Qdrant backup snapshots
- One inference VM with cloud-init that starts Docker Compose services: Ollama (LLM), Qdrant (vector store), and a Python gateway on port 8080
- Gateway routes for
/transcribe,/summarize, and/indexcreator workflows - 0 floating IPs by default (access via bastion or VPN to the private gateway URL)
Enable enable_public_api to attach one floating IP when an external client must reach the gateway without a bastion.
Prerequisites#
- OpenTofu or Terraform >= 1.6.0
- OpenStack application credentials (
source openrc.sh) - EC2-compatible Object Storage credentials. See S3 storage ACL template.
- One floating IP available when enabling public API access
Parameters#
| Parameter | Description | Default |
|---|---|---|
key_name | SSH keypair name (must already exist in your project) | required |
data_bucket_name | Bucket for inputs, outputs, and snapshots | required |
s3_access_key / s3_secret_key | EC2-compat Object Storage credentials | required |
inference_flavor | CPU flavor for Ollama and whisper workloads | c2a.large |
ollama_model | Model tag pulled at first boot | llama3.2:1b |
enable_public_api | Attach one FIP and open gateway port 8080 | false |
api_token | Optional bearer token for gateway auth | "" |
image_name | Boot image name | Ubuntu-24.04 |
external_network | Shared external network for router gateway and floating IP | PublicStatic |
private_cidr | Private subnet CIDR for the inference instance | 192.168.85.0/24 |
instance_name | Inference instance display name | creator-ai-worker |
admin_cidr | Source CIDR allowed for SSH and API access | 0.0.0.0/0 |
s3_endpoint | Quake AI S3-compatible endpoint URL | https://object.us-east-1.rumble.cloud |
s3_region | S3 region identifier for the AWS provider | us-east-1 |
Cost and sizing#
Size inference_flavor for the Ollama model you pull and any concurrent whisper transcription. CPU inference suits small models (roughly 7B-8B parameter ceiling on typical flavors). Object Storage charges follow your plan allowance for stored media and Qdrant snapshots.
When to use this pattern#
Wire transcribe-to-summarize-to-index creator workflows (show notes, captions, search) on a single stack without a managed AI SaaS. For batch audio loudness only, see the audio post-production worker template.
Estimated cost#
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Inference
c2a.large · 2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
c2a.large
2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
2 vCPU + 4 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (80 GiB)
80 GiB at $0.08/GiB/mo
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Object storage (usage-based)
Object storage
1 bucket. The first 1 TB is included, then $10.00 per TB each month. You pay for what you store, so this line depends on usage.
Assumes: 2 TB stored is $10/mo; 5 TB stored is $40/mo. Within the included allotment it stays $0.
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn moreObject storage upload and download
No separate charges for uploading or downloading object storage data.
Most object storage providers meter egress and API requests separately from stored capacity.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Configure your estimate
Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.
Starting template
The required baseline, always included.
Pick how much you expect to store to fold it into the total.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on dedicated CPU
The headline estimate above; predictable steady-load performance.
Saves $45.50/mo while you build on shared CPU.
Shared flavors carry less RAM (c2a.large (4 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Template source#
Show source (7 files)Hide source
data "openstack_images_image_v2" "os" {
name = var.image_name
most_recent = true
}
data "openstack_networking_network_v2" "external" {
name = var.external_network
}
resource "aws_s3_bucket" "data" {
bucket = var.data_bucket_name
}
resource "aws_s3_bucket_acl" "data" {
bucket = aws_s3_bucket.data.id
acl = "private"
}
resource "openstack_networking_network_v2" "private" {
name = "${var.instance_name}-private"
admin_state_up = true
}
resource "openstack_networking_subnet_v2" "private" {
name = "${var.instance_name}-private-sn"
network_id = openstack_networking_network_v2.private.id
cidr = var.private_cidr
ip_version = 4
enable_dhcp = true
dns_nameservers = ["8.8.8.8", "8.8.4.4"]
}
resource "openstack_networking_router_v2" "router" {
name = "${var.instance_name}-router"
external_network_id = data.openstack_networking_network_v2.external.id
admin_state_up = true
}
resource "openstack_networking_router_interface_v2" "private" {
router_id = openstack_networking_router_v2.router.id
subnet_id = openstack_networking_subnet_v2.private.id
}
resource "openstack_networking_secgroup_v2" "inference" {
name = "${var.instance_name}-sg"
description = "Creator-AI gateway API, SSH from admin CIDR"
}
resource "openstack_networking_secgroup_rule_v2" "ssh" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 22
port_range_max = 22
remote_ip_prefix = var.admin_cidr
security_group_id = openstack_networking_secgroup_v2.inference.id
}
resource "openstack_networking_secgroup_rule_v2" "gateway_api" {
count = var.enable_public_api ? 1 : 0
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 8080
port_range_max = 8080
remote_ip_prefix = var.admin_cidr
security_group_id = openstack_networking_secgroup_v2.inference.id
}
resource "openstack_networking_port_v2" "inference" {
name = "${var.instance_name}-port"
network_id = openstack_networking_network_v2.private.id
security_group_ids = [openstack_networking_secgroup_v2.inference.id]
fixed_ip {
subnet_id = openstack_networking_subnet_v2.private.id
}
depends_on = [openstack_networking_router_interface_v2.private]
}
resource "openstack_compute_instance_v2" "inference" {
name = var.instance_name
flavor_name = var.inference_flavor
key_pair = var.key_name
user_data = templatefile("${path.module}/cloud-init/inference.yaml", {
s3_endpoint = var.s3_endpoint
s3_region = var.s3_region
s3_access_key = var.s3_access_key
s3_secret_key = var.s3_secret_key
data_bucket = aws_s3_bucket.data.id
ollama_model = var.ollama_model
api_token = var.api_token
})
block_device {
uuid = data.openstack_images_image_v2.os.id
source_type = "image"
destination_type = "volume"
volume_size = 80
boot_index = 0
delete_on_termination = true
}
network {
port = openstack_networking_port_v2.inference.id
}
depends_on = [openstack_networking_router_interface_v2.private]
}
resource "openstack_networking_floatingip_v2" "inference" {
count = var.enable_public_api ? 1 : 0
pool = var.external_network
}
resource "openstack_networking_floatingip_associate_v2" "inference" {
count = var.enable_public_api ? 1 : 0
floating_ip = openstack_networking_floatingip_v2.inference[0].address
port_id = openstack_networking_port_v2.inference.id
}
variable "key_name" {
description = "Existing SSH keypair name in the project for compute instances"
type = string
}
variable "s3_access_key" {
description = "EC2-compatible access key for Object Storage (Keystone ec2 credentials create)"
type = string
sensitive = true
}
variable "s3_secret_key" {
description = "EC2-compatible secret key paired with s3_access_key"
type = string
sensitive = true
}
variable "data_bucket_name" {
description = "S3 bucket name for input media, generated captions, and Qdrant backup snapshots"
type = string
}
variable "inference_flavor" {
description = "Compute flavor for the inference VM (CPU-only; size for concurrent whisper/Ollama workloads)"
type = string
default = "c2a.large"
}
variable "image_name" {
description = "Boot image name"
type = string
default = "Ubuntu-24.04"
}
variable "external_network" {
description = "Shared external network for router gateway and floating IPs; defaults to PublicStatic (persisted FIP / production pattern). Override with PublicEphemeral for ephemeral demos."
type = string
default = "PublicStatic"
}
variable "private_cidr" {
description = "Private subnet CIDR for the inference instance"
type = string
default = "192.168.85.0/24"
}
variable "instance_name" {
description = "Inference instance display name"
type = string
default = "creator-ai-worker"
}
variable "ollama_model" {
description = "Default Ollama model tag pulled at first boot (keep small for CPU; for example llama3.2:1b)"
type = string
default = "llama3.2:1b"
}
variable "enable_public_api" {
description = "When true, attach one floating IP and expose the gateway API on port 8080"
type = bool
default = false
}
variable "admin_cidr" {
description = "Source CIDR allowed for SSH (22/tcp) and API access when enable_public_api is true"
type = string
default = "0.0.0.0/0"
}
variable "api_token" {
description = "Bearer token clients pass to the gateway API (empty generates no auth check in the stub gateway)"
type = string
default = ""
sensitive = true
}
variable "s3_endpoint" {
description = "Quake AI S3-compatible endpoint URL"
type = string
default = "https://object.us-east-1.rumble.cloud"
}
variable "s3_region" {
description = "S3 region identifier passed to the AWS provider"
type = string
default = "us-east-1"
}
output "inference_instance_id" {
description = "Nova instance ID for the Creator-AI inference stack"
value = openstack_compute_instance_v2.inference.id
}
output "inference_private_ip" {
description = "Private IPv4 address for bastion or VPN access to the gateway API"
value = openstack_networking_port_v2.inference.all_fixed_ips[0]
}
output "inference_floating_ip" {
description = "Public floating IP when enable_public_api is true"
value = var.enable_public_api ? openstack_networking_floatingip_v2.inference[0].address : ""
}
output "gateway_url" {
description = "HTTP gateway base URL for transcribe/summarize/RAG workflows"
value = var.enable_public_api ? "http://${openstack_networking_floatingip_v2.inference[0].address}:8080" : "http://${openstack_networking_port_v2.inference.all_fixed_ips[0]}:8080"
}
output "data_bucket" {
description = "Object Storage bucket for inputs, outputs, and Qdrant snapshots"
value = aws_s3_bucket.data.id
}
output "floating_ip_count" {
description = "Floating IPs this template allocates (zero by default; one when enable_public_api is true)"
value = var.enable_public_api ? 1 : 0
}
terraform {
required_version = ">= 1.6.0"
required_providers {
openstack = {
source = "terraform-provider-openstack/openstack"
version = "~> 2.0"
}
aws = {
source = "hashicorp/aws"
version = "~> 5.0"
}
}
}
provider "openstack" {}
provider "aws" {
region = var.s3_region
access_key = var.s3_access_key
secret_key = var.s3_secret_key
skip_credentials_validation = true
skip_metadata_api_check = true
skip_requesting_account_id = true
endpoints {
s3 = var.s3_endpoint
}
}
# Required
key_name = "YOUR_KEY_NAME"
data_bucket_name = "my-creator-ai-data"
s3_access_key = "YOUR_EC2_ACCESS_KEY"
s3_secret_key = "YOUR_EC2_SECRET_KEY"
# inference_flavor = "c2a.large"
# ollama_model = "llama3.2:1b"
# enable_public_api = false
# api_token = "your-gateway-token"
#cloud-config
package_update: true
write_files:
- path: /etc/creator-ai.env
owner: root:root
permissions: "0600"
content: |
AWS_ACCESS_KEY_ID=${s3_access_key}
AWS_SECRET_ACCESS_KEY=${s3_secret_key}
AWS_DEFAULT_REGION=${s3_region}
AWS_ENDPOINT_URL_S3=${s3_endpoint}
DATA_BUCKET=${data_bucket}
OLLAMA_MODEL=${ollama_model}
API_TOKEN=${api_token}
- path: /opt/creator-ai/docker-compose.yml
owner: root:root
permissions: "0644"
content: |
services:
ollama:
image: ollama/ollama:latest
restart: unless-stopped
volumes:
- ollama_data:/root/.ollama
ports:
- "11434:11434"
qdrant:
image: qdrant/qdrant:v1.12.5
restart: unless-stopped
volumes:
- qdrant_data:/qdrant/storage
ports:
- "6333:6333"
gateway:
image: python:3.12-slim
restart: unless-stopped
network_mode: host
volumes:
- /opt/creator-ai/gateway.py:/app/gateway.py:ro
- /etc/creator-ai.env:/etc/creator-ai.env:ro
command: ["python3", "/app/gateway.py"]
volumes:
ollama_data:
qdrant_data:
- path: /opt/creator-ai/gateway.py
owner: root:root
permissions: "0755"
content: |
#!/usr/bin/env python3
import json, os, subprocess, urllib.request
from http.server import BaseHTTPRequestHandler, HTTPServer
def load_env(path="/etc/creator-ai.env"):
env = {}
for line in open(path):
line = line.strip()
if line and not line.startswith("#") and "=" in line:
k, v = line.split("=", 1)
env[k] = v
return env
ENV = load_env()
OLLAMA = "http://127.0.0.1:11434"
QDRANT = "http://127.0.0.1:6333"
class Handler(BaseHTTPRequestHandler):
def _json(self, code, body):
data = json.dumps(body).encode()
self.send_response(code)
self.send_header("Content-Type", "application/json")
self.send_header("Content-Length", str(len(data)))
self.end_headers()
self.wfile.write(data)
def _authorized(self):
token = ENV.get("API_TOKEN", "")
if not token:
return True
auth = self.headers.get("Authorization", "")
return auth == f"Bearer {token}"
def do_GET(self):
if self.path == "/health":
return self._json(200, {"status": "ok"})
if not self._authorized():
return self._json(401, {"error": "unauthorized"})
return self._json(404, {"error": "not found"})
def do_POST(self):
if not self._authorized():
return self._json(401, {"error": "unauthorized"})
length = int(self.headers.get("Content-Length", 0))
body = json.loads(self.rfile.read(length).decode() if length else "{}")
if self.path == "/transcribe":
key = body.get("input_key", "")
if not key:
return self._json(400, {"error": "input_key required"})
return self._json(202, {"status": "queued", "input_key": key, "engine": "whisper.cpp"})
if self.path == "/summarize":
text = body.get("text", "")
model = ENV.get("OLLAMA_MODEL", "llama3.2:1b")
payload = json.dumps({"model": model, "prompt": f"Summarize for show notes:\n{text}", "stream": False}).encode()
req = urllib.request.Request(f"{OLLAMA}/api/generate", data=payload, headers={"Content-Type": "application/json"})
with urllib.request.urlopen(req, timeout=120) as resp:
out = json.loads(resp.read().decode())
return self._json(200, {"summary": out.get("response", "")})
if self.path == "/index":
doc_id = body.get("doc_id", "")
text = body.get("text", "")
if not doc_id or not text:
return self._json(400, {"error": "doc_id and text required"})
point = {"points": [{"id": doc_id, "vector": [0.0] * 384, "payload": {"text": text}}]}
req = urllib.request.Request(f"{QDRANT}/collections/creator/points", json.dumps(point).encode(), {"Content-Type": "application/json"}, method="PUT")
try:
urllib.request.urlopen(req, timeout=30)
except Exception:
coll = json.dumps({"vectors": {"size": 384, "distance": "Cosine"}}).encode()
urllib.request.urlopen(urllib.request.Request(f"{QDRANT}/collections/creator", coll, {"Content-Type": "application/json"}, method="PUT"), timeout=30)
urllib.request.urlopen(req, timeout=30)
return self._json(200, {"indexed": doc_id})
return self._json(404, {"error": "not found"})
def log_message(self, fmt, *args):
return
if __name__ == "__main__":
HTTPServer(("0.0.0.0", 8080), Handler).serve_forever()
runcmd:
- |
set -e
export DEBIAN_FRONTEND=noninteractive
apt-get update
apt-get install -y docker.io docker-compose-v2 curl python3-pip
pip3 install --break-system-packages awscli
systemctl enable --now docker
docker compose -f /opt/creator-ai/docker-compose.yml up -d
docker exec creator-ai-ollama-1 ollama pull ${ollama_model} || \
docker compose -f /opt/creator-ai/docker-compose.yml exec -T ollama ollama pull ${ollama_model} || true
# Creator-AI inference worker
Single inference VM running Ollama, Qdrant, and a lightweight gateway API for creator workflows (transcribe, summarize, index). CPU inference by default.
**Network class:** production — `external_network` defaults to `PublicStatic` for persisted floating IPs and multi-tier stacks; override with `PublicEphemeral` for ephemeral demos.
## Prerequisites
- OpenTofu >= 1.6.0 or Terraform >= 1.6.0
- Quake AI account with OpenStack credentials
- EC2-compatible Object Storage credentials
- Optional: one floating IP quota slot when `enable_public_api` is true
## Usage
1. Clone or copy this template directory
2. Copy `terraform.tfvars.example` to `terraform.tfvars` and fill in your values
3. Source your OpenStack credentials: `source openrc.sh`
4. Initialize: `tofu init`
5. Preview: `tofu plan`
6. Apply: `tofu apply`
POST to `/transcribe`, `/summarize`, and `/index` on the gateway URL. Upload source media to the data bucket first.
## Variables
| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `key_name` | string | yes | n/a | Existing SSH keypair name in the project for compute instances |
| `data_bucket_name` | string | yes | n/a | Inputs, outputs, Qdrant snapshots |
| `inference_flavor` | string | no | `c2a.large` | CPU flavor for inference |
| `ollama_model` | string | no | `llama3.2:1b` | Default Ollama model (keep small on CPU) |
| `enable_public_api` | bool | no | `false` | Attach one FIP for external API access |
## Documentation
Full documentation: [Creator-AI inference worker template](/docs/automation/templates/creator-ai-worker)
## Honesty note
CPU inference is the default; latency and model size are bounded versus GPU. See the self-hosted AI concept page for platform limits. No managed AI service is implied.
Resources, parameters, and variables
key_namerequireds3_access_keyrequireds3_secret_keyrequireddata_bucket_namerequiredinference_flavor="c2a.large"image_name="Ubuntu-24.04"external_network="PublicStatic"private_cidr="192.168.85.0/24"instance_name="creator-ai-worker"ollama_model="llama3.2:1b"enable_public_api=falseadmin_cidr="0.0.0.0/0"api_token=""s3_endpoint="https://object.us-east-1.rumble.cloud"s3_region="us-east-1"
Outputs#
| Output | Description |
|---|---|
inference_private_ip | Private IP for bastion/VPN access |
inference_floating_ip | Public IP when enable_public_api is true |
gateway_url | Base URL for /transcribe, /summarize, /index |
data_bucket | Object Storage bucket name |
floating_ip_count | 0 by default; 1 when public API is enabled |
See also#
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.
Last validated: 07.07.2026
See Also
Deploy the Creator-AI inference worker template with OpenTofu
Shares: Self Hosted Ai, Cloud Init
Compute concepts
Shares: Self Hosted Ai, Cloud Init
Deploy the audio post-production worker template with OpenTofu
Shares: Cloud Init, Object Storage
Deploy the live RTMP/SRT ingest and restream template with OpenTofu
Shares: Cloud Init, Object Storage
Deploy the CPU render-farm worker pool template with OpenTofu
Shares: Cloud Init, Object Storage