Skip to content
IaC Templates

Creator-AI inference worker

Template · Updated Jul 2026
Validated Jul 2026

Creator-AI inference worker

This pattern composes Compute and Object Storage.

What this template deploys#

  • Private network and router for the inference instance
  • One Object Storage bucket for input media, generated captions, and Qdrant backup snapshots
  • One inference VM with cloud-init that starts Docker Compose services: Ollama (LLM), Qdrant (vector store), and a Python gateway on port 8080
  • Gateway routes for /transcribe, /summarize, and /index creator workflows
  • 0 floating IPs by default (access via bastion or VPN to the private gateway URL)

Enable enable_public_api to attach one floating IP when an external client must reach the gateway without a bastion.

Prerequisites#

  • OpenTofu or Terraform >= 1.6.0
  • OpenStack application credentials (source openrc.sh)
  • EC2-compatible Object Storage credentials. See S3 storage ACL template.
  • One floating IP available when enabling public API access

Parameters#

ParameterDescriptionDefault
key_nameSSH keypair name (must already exist in your project)required
data_bucket_nameBucket for inputs, outputs, and snapshotsrequired
s3_access_key / s3_secret_keyEC2-compat Object Storage credentialsrequired
inference_flavorCPU flavor for Ollama and whisper workloadsc2a.large
ollama_modelModel tag pulled at first bootllama3.2:1b
enable_public_apiAttach one FIP and open gateway port 8080false
api_tokenOptional bearer token for gateway auth""
image_nameBoot image nameUbuntu-24.04
external_networkShared external network for router gateway and floating IPPublicStatic
private_cidrPrivate subnet CIDR for the inference instance192.168.85.0/24
instance_nameInference instance display namecreator-ai-worker
admin_cidrSource CIDR allowed for SSH and API access0.0.0.0/0
s3_endpointQuake AI S3-compatible endpoint URLhttps://object.us-east-1.rumble.cloud
s3_regionS3 region identifier for the AWS providerus-east-1

Cost and sizing#

Size inference_flavor for the Ollama model you pull and any concurrent whisper transcription. CPU inference suits small models (roughly 7B-8B parameter ceiling on typical flavors). Object Storage charges follow your plan allowance for stored media and Qdrant snapshots.

When to use this pattern#

Wire transcribe-to-summarize-to-index creator workflows (show notes, captions, search) on a single stack without a managed AI SaaS. For batch audio loudness only, see the audio post-production worker template.

Estimated cost#

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on dedicated vCPU.

Starting template$63.40/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Inference

c2a.large · 2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps

$62.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

c2a.large

2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps

$62.00

Compute + RAM rate basis

2 vCPU + 4 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (80 GiB)

80 GiB at $0.08/GiB/mo

$6.40

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Object storage (usage-based)

Object storage

1 bucket. The first 1 TB is included, then $10.00 per TB each month. You pay for what you store, so this line depends on usage.

$0–$40/mo

Assumes: 2 TB stored is $10/mo; 5 TB stored is $40/mo. Within the included allotment it stays $0.

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Object storage upload and download

No separate charges for uploading or downloading object storage data.

Most object storage providers meter egress and API requests separately from stored capacity.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Configure your estimate

Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.

Starting template

The required baseline, always included.

$63.40/mo

Pick how much you expect to store to fold it into the total.

$0.00/mo
Your configured estimate$63.40/mo

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$17.90/mo

Production on dedicated CPU

The headline estimate above; predictable steady-load performance.

$63.40/mo

Saves $45.50/mo while you build on shared CPU.

Shared flavors carry less RAM (c2a.large (4 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Template source#

7 files. Download the zip or expand to copy any file.Download creator-ai-worker.zip
Show source (7 files)
main.tfHCL
data "openstack_images_image_v2" "os" {
  name        = var.image_name
  most_recent = true
}

data "openstack_networking_network_v2" "external" {
  name = var.external_network
}

resource "aws_s3_bucket" "data" {
  bucket = var.data_bucket_name
}

resource "aws_s3_bucket_acl" "data" {
  bucket = aws_s3_bucket.data.id
  acl    = "private"
}


resource "openstack_networking_network_v2" "private" {
  name           = "${var.instance_name}-private"
  admin_state_up = true
}

resource "openstack_networking_subnet_v2" "private" {
  name            = "${var.instance_name}-private-sn"
  network_id      = openstack_networking_network_v2.private.id
  cidr            = var.private_cidr
  ip_version      = 4
  enable_dhcp     = true
  dns_nameservers = ["8.8.8.8", "8.8.4.4"]
}

resource "openstack_networking_router_v2" "router" {
  name                = "${var.instance_name}-router"
  external_network_id = data.openstack_networking_network_v2.external.id
  admin_state_up      = true
}

resource "openstack_networking_router_interface_v2" "private" {
  router_id = openstack_networking_router_v2.router.id
  subnet_id = openstack_networking_subnet_v2.private.id
}

resource "openstack_networking_secgroup_v2" "inference" {
  name        = "${var.instance_name}-sg"
  description = "Creator-AI gateway API, SSH from admin CIDR"
}

resource "openstack_networking_secgroup_rule_v2" "ssh" {
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 22
  port_range_max    = 22
  remote_ip_prefix  = var.admin_cidr
  security_group_id = openstack_networking_secgroup_v2.inference.id
}

resource "openstack_networking_secgroup_rule_v2" "gateway_api" {
  count             = var.enable_public_api ? 1 : 0
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 8080
  port_range_max    = 8080
  remote_ip_prefix  = var.admin_cidr
  security_group_id = openstack_networking_secgroup_v2.inference.id
}

resource "openstack_networking_port_v2" "inference" {
  name               = "${var.instance_name}-port"
  network_id         = openstack_networking_network_v2.private.id
  security_group_ids = [openstack_networking_secgroup_v2.inference.id]

  fixed_ip {
    subnet_id = openstack_networking_subnet_v2.private.id
  }

  depends_on = [openstack_networking_router_interface_v2.private]
}

resource "openstack_compute_instance_v2" "inference" {
  name        = var.instance_name
  flavor_name = var.inference_flavor
  key_pair    = var.key_name

  user_data = templatefile("${path.module}/cloud-init/inference.yaml", {
    s3_endpoint   = var.s3_endpoint
    s3_region     = var.s3_region
    s3_access_key = var.s3_access_key
    s3_secret_key = var.s3_secret_key
    data_bucket   = aws_s3_bucket.data.id
    ollama_model  = var.ollama_model
    api_token     = var.api_token
  })

  block_device {
    uuid                  = data.openstack_images_image_v2.os.id
    source_type           = "image"
    destination_type      = "volume"
    volume_size           = 80
    boot_index            = 0
    delete_on_termination = true
  }

  network {
    port = openstack_networking_port_v2.inference.id
  }

  depends_on = [openstack_networking_router_interface_v2.private]
}

resource "openstack_networking_floatingip_v2" "inference" {
  count = var.enable_public_api ? 1 : 0
  pool  = var.external_network
}

resource "openstack_networking_floatingip_associate_v2" "inference" {
  count       = var.enable_public_api ? 1 : 0
  floating_ip = openstack_networking_floatingip_v2.inference[0].address
  port_id     = openstack_networking_port_v2.inference.id
}
variables.tfHCL
variable "key_name" {
  description = "Existing SSH keypair name in the project for compute instances"
  type        = string
}

variable "s3_access_key" {
  description = "EC2-compatible access key for Object Storage (Keystone ec2 credentials create)"
  type        = string
  sensitive   = true
}

variable "s3_secret_key" {
  description = "EC2-compatible secret key paired with s3_access_key"
  type        = string
  sensitive   = true
}

variable "data_bucket_name" {
  description = "S3 bucket name for input media, generated captions, and Qdrant backup snapshots"
  type        = string
}

variable "inference_flavor" {
  description = "Compute flavor for the inference VM (CPU-only; size for concurrent whisper/Ollama workloads)"
  type        = string
  default     = "c2a.large"
}

variable "image_name" {
  description = "Boot image name"
  type        = string
  default     = "Ubuntu-24.04"
}

variable "external_network" {
  description = "Shared external network for router gateway and floating IPs; defaults to PublicStatic (persisted FIP / production pattern). Override with PublicEphemeral for ephemeral demos."
  type        = string
  default     = "PublicStatic"
}

variable "private_cidr" {
  description = "Private subnet CIDR for the inference instance"
  type        = string
  default     = "192.168.85.0/24"
}

variable "instance_name" {
  description = "Inference instance display name"
  type        = string
  default     = "creator-ai-worker"
}

variable "ollama_model" {
  description = "Default Ollama model tag pulled at first boot (keep small for CPU; for example llama3.2:1b)"
  type        = string
  default     = "llama3.2:1b"
}

variable "enable_public_api" {
  description = "When true, attach one floating IP and expose the gateway API on port 8080"
  type        = bool
  default     = false
}

variable "admin_cidr" {
  description = "Source CIDR allowed for SSH (22/tcp) and API access when enable_public_api is true"
  type        = string
  default     = "0.0.0.0/0"
}

variable "api_token" {
  description = "Bearer token clients pass to the gateway API (empty generates no auth check in the stub gateway)"
  type        = string
  default     = ""
  sensitive   = true
}

variable "s3_endpoint" {
  description = "Quake AI S3-compatible endpoint URL"
  type        = string
  default     = "https://object.us-east-1.rumble.cloud"
}

variable "s3_region" {
  description = "S3 region identifier passed to the AWS provider"
  type        = string
  default     = "us-east-1"
}
outputs.tfHCL
output "inference_instance_id" {
  description = "Nova instance ID for the Creator-AI inference stack"
  value       = openstack_compute_instance_v2.inference.id
}

output "inference_private_ip" {
  description = "Private IPv4 address for bastion or VPN access to the gateway API"
  value       = openstack_networking_port_v2.inference.all_fixed_ips[0]
}

output "inference_floating_ip" {
  description = "Public floating IP when enable_public_api is true"
  value       = var.enable_public_api ? openstack_networking_floatingip_v2.inference[0].address : ""
}

output "gateway_url" {
  description = "HTTP gateway base URL for transcribe/summarize/RAG workflows"
  value       = var.enable_public_api ? "http://${openstack_networking_floatingip_v2.inference[0].address}:8080" : "http://${openstack_networking_port_v2.inference.all_fixed_ips[0]}:8080"
}

output "data_bucket" {
  description = "Object Storage bucket for inputs, outputs, and Qdrant snapshots"
  value       = aws_s3_bucket.data.id
}

output "floating_ip_count" {
  description = "Floating IPs this template allocates (zero by default; one when enable_public_api is true)"
  value       = var.enable_public_api ? 1 : 0
}
versions.tfHCL
terraform {
  required_version = ">= 1.6.0"

  required_providers {
    openstack = {
      source  = "terraform-provider-openstack/openstack"
      version = "~> 2.0"
    }
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.0"
    }
  }
}

provider "openstack" {}

provider "aws" {
  region                      = var.s3_region
  access_key                  = var.s3_access_key
  secret_key                  = var.s3_secret_key
  skip_credentials_validation = true
  skip_metadata_api_check     = true
  skip_requesting_account_id  = true

  endpoints {
    s3 = var.s3_endpoint
  }
}
terraform.tfvars.exampleHCL
# Required
key_name          = "YOUR_KEY_NAME"
data_bucket_name  = "my-creator-ai-data"
s3_access_key     = "YOUR_EC2_ACCESS_KEY"
s3_secret_key     = "YOUR_EC2_SECRET_KEY"

# inference_flavor = "c2a.large"
# ollama_model = "llama3.2:1b"
# enable_public_api = false
# api_token = "your-gateway-token"
cloud-init/inference.yamlYAML
#cloud-config
package_update: true
write_files:
  - path: /etc/creator-ai.env
    owner: root:root
    permissions: "0600"
    content: |
      AWS_ACCESS_KEY_ID=${s3_access_key}
      AWS_SECRET_ACCESS_KEY=${s3_secret_key}
      AWS_DEFAULT_REGION=${s3_region}
      AWS_ENDPOINT_URL_S3=${s3_endpoint}
      DATA_BUCKET=${data_bucket}
      OLLAMA_MODEL=${ollama_model}
      API_TOKEN=${api_token}
  - path: /opt/creator-ai/docker-compose.yml
    owner: root:root
    permissions: "0644"
    content: |
      services:
        ollama:
          image: ollama/ollama:latest
          restart: unless-stopped
          volumes:
            - ollama_data:/root/.ollama
          ports:
            - "11434:11434"
        qdrant:
          image: qdrant/qdrant:v1.12.5
          restart: unless-stopped
          volumes:
            - qdrant_data:/qdrant/storage
          ports:
            - "6333:6333"
        gateway:
          image: python:3.12-slim
          restart: unless-stopped
          network_mode: host
          volumes:
            - /opt/creator-ai/gateway.py:/app/gateway.py:ro
            - /etc/creator-ai.env:/etc/creator-ai.env:ro
          command: ["python3", "/app/gateway.py"]
      volumes:
        ollama_data:
        qdrant_data:
  - path: /opt/creator-ai/gateway.py
    owner: root:root
    permissions: "0755"
    content: |
      #!/usr/bin/env python3
      import json, os, subprocess, urllib.request
      from http.server import BaseHTTPRequestHandler, HTTPServer

      def load_env(path="/etc/creator-ai.env"):
          env = {}
          for line in open(path):
              line = line.strip()
              if line and not line.startswith("#") and "=" in line:
                  k, v = line.split("=", 1)
                  env[k] = v
          return env

      ENV = load_env()
      OLLAMA = "http://127.0.0.1:11434"
      QDRANT = "http://127.0.0.1:6333"

      class Handler(BaseHTTPRequestHandler):
          def _json(self, code, body):
              data = json.dumps(body).encode()
              self.send_response(code)
              self.send_header("Content-Type", "application/json")
              self.send_header("Content-Length", str(len(data)))
              self.end_headers()
              self.wfile.write(data)

          def _authorized(self):
              token = ENV.get("API_TOKEN", "")
              if not token:
                  return True
              auth = self.headers.get("Authorization", "")
              return auth == f"Bearer {token}"

          def do_GET(self):
              if self.path == "/health":
                  return self._json(200, {"status": "ok"})
              if not self._authorized():
                  return self._json(401, {"error": "unauthorized"})
              return self._json(404, {"error": "not found"})

          def do_POST(self):
              if not self._authorized():
                  return self._json(401, {"error": "unauthorized"})
              length = int(self.headers.get("Content-Length", 0))
              body = json.loads(self.rfile.read(length).decode() if length else "{}")
              if self.path == "/transcribe":
                  key = body.get("input_key", "")
                  if not key:
                      return self._json(400, {"error": "input_key required"})
                  return self._json(202, {"status": "queued", "input_key": key, "engine": "whisper.cpp"})
              if self.path == "/summarize":
                  text = body.get("text", "")
                  model = ENV.get("OLLAMA_MODEL", "llama3.2:1b")
                  payload = json.dumps({"model": model, "prompt": f"Summarize for show notes:\n{text}", "stream": False}).encode()
                  req = urllib.request.Request(f"{OLLAMA}/api/generate", data=payload, headers={"Content-Type": "application/json"})
                  with urllib.request.urlopen(req, timeout=120) as resp:
                      out = json.loads(resp.read().decode())
                  return self._json(200, {"summary": out.get("response", "")})
              if self.path == "/index":
                  doc_id = body.get("doc_id", "")
                  text = body.get("text", "")
                  if not doc_id or not text:
                      return self._json(400, {"error": "doc_id and text required"})
                  point = {"points": [{"id": doc_id, "vector": [0.0] * 384, "payload": {"text": text}}]}
                  req = urllib.request.Request(f"{QDRANT}/collections/creator/points", json.dumps(point).encode(), {"Content-Type": "application/json"}, method="PUT")
                  try:
                      urllib.request.urlopen(req, timeout=30)
                  except Exception:
                      coll = json.dumps({"vectors": {"size": 384, "distance": "Cosine"}}).encode()
                      urllib.request.urlopen(urllib.request.Request(f"{QDRANT}/collections/creator", coll, {"Content-Type": "application/json"}, method="PUT"), timeout=30)
                      urllib.request.urlopen(req, timeout=30)
                  return self._json(200, {"indexed": doc_id})
              return self._json(404, {"error": "not found"})

          def log_message(self, fmt, *args):
              return

      if __name__ == "__main__":
          HTTPServer(("0.0.0.0", 8080), Handler).serve_forever()
runcmd:
  - |
    set -e
    export DEBIAN_FRONTEND=noninteractive
    apt-get update
    apt-get install -y docker.io docker-compose-v2 curl python3-pip
    pip3 install --break-system-packages awscli
    systemctl enable --now docker
    docker compose -f /opt/creator-ai/docker-compose.yml up -d
    docker exec creator-ai-ollama-1 ollama pull ${ollama_model} || \
      docker compose -f /opt/creator-ai/docker-compose.yml exec -T ollama ollama pull ${ollama_model} || true
README.mdMarkdown
# Creator-AI inference worker

Single inference VM running Ollama, Qdrant, and a lightweight gateway API for creator workflows (transcribe, summarize, index). CPU inference by default.


**Network class:** production — `external_network` defaults to `PublicStatic` for persisted floating IPs and multi-tier stacks; override with `PublicEphemeral` for ephemeral demos.

## Prerequisites

- OpenTofu >= 1.6.0 or Terraform >= 1.6.0
- Quake AI account with OpenStack credentials
- EC2-compatible Object Storage credentials
- Optional: one floating IP quota slot when `enable_public_api` is true

## Usage

1. Clone or copy this template directory
2. Copy `terraform.tfvars.example` to `terraform.tfvars` and fill in your values
3. Source your OpenStack credentials: `source openrc.sh`
4. Initialize: `tofu init`
5. Preview: `tofu plan`
6. Apply: `tofu apply`

POST to `/transcribe`, `/summarize`, and `/index` on the gateway URL. Upload source media to the data bucket first.

## Variables

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `key_name` | string | yes | n/a | Existing SSH keypair name in the project for compute instances |
| `data_bucket_name` | string | yes | n/a | Inputs, outputs, Qdrant snapshots |
| `inference_flavor` | string | no | `c2a.large` | CPU flavor for inference |
| `ollama_model` | string | no | `llama3.2:1b` | Default Ollama model (keep small on CPU) |
| `enable_public_api` | bool | no | `false` | Attach one FIP for external API access |

## Documentation

Full documentation: [Creator-AI inference worker template](/docs/automation/templates/creator-ai-worker)

## Honesty note

CPU inference is the default; latency and model size are bounded versus GPU. See the self-hosted AI concept page for platform limits. No managed AI service is implied.
Resources, parameters, and variables
Provisions
Parameterized by
Variables
  • key_namerequired
  • s3_access_keyrequired
  • s3_secret_keyrequired
  • data_bucket_namerequired
  • inference_flavor="c2a.large"
  • image_name="Ubuntu-24.04"
  • external_network="PublicStatic"
  • private_cidr="192.168.85.0/24"
  • instance_name="creator-ai-worker"
  • ollama_model="llama3.2:1b"
  • enable_public_api=false
  • admin_cidr="0.0.0.0/0"
  • api_token=""
  • s3_endpoint="https://object.us-east-1.rumble.cloud"
  • s3_region="us-east-1"

Outputs#

OutputDescription
inference_private_ipPrivate IP for bastion/VPN access
inference_floating_ipPublic IP when enable_public_api is true
gateway_urlBase URL for /transcribe, /summarize, /index
data_bucketObject Storage bucket name
floating_ip_count0 by default; 1 when public API is enabled

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

For the full policy, see Usage Guidelines.

Last validated: 07.07.2026

Was this page helpful?