# Creator-AI inference worker

Source: https://docs.quake.ai/resources/iac-templates/creator-ai-worker
Markdown: https://docs.quake.ai/resources/iac-templates/creator-ai-worker.md

---

# Creator-AI inference worker

This pattern composes Compute and Object Storage.



Quake AI does not provide managed AI inference. This template provisions a customer-owned VM running Ollama, Qdrant, whisper.cpp, and Piper as containers on CPU. Performance is bounded versus GPU: keep Ollama models small and expect higher latency. See [Self-Hosted AI on Cloud Infrastructure](/docs/compute/concepts/self-hosted-ai) for the platform CPU-only ceiling.



## What this template deploys

- Private network and router for the inference instance
- One Object Storage bucket for input media, generated captions, and Qdrant backup snapshots
- One inference VM with cloud-init that starts Docker Compose services: Ollama (LLM), Qdrant (vector store), and a Python gateway on port 8080
- Gateway routes for `/transcribe`, `/summarize`, and `/index` creator workflows
- **0 floating IPs** by default (access via bastion or VPN to the private gateway URL)

Enable `enable_public_api` to attach one floating IP when an external client must reach the gateway without a bastion.

## Prerequisites

- OpenTofu or Terraform >= 1.6.0
- OpenStack application credentials (`source openrc.sh`)
- EC2-compatible Object Storage credentials. See [S3 storage ACL template](/resources/iac-templates/s3-storage-acl).
- One floating IP available when enabling public API access

## Parameters

| Parameter | Description | Default |
| --- | --- | --- |
| `key_name` | SSH keypair name (must already exist in your project) | required |
| `data_bucket_name` | Bucket for inputs, outputs, and snapshots | required |
| `s3_access_key` / `s3_secret_key` | EC2-compat Object Storage credentials | required |
| `inference_flavor` | CPU flavor for Ollama and whisper workloads | `c2a.large` |
| `ollama_model` | Model tag pulled at first boot | `llama3.2:1b` |
| `enable_public_api` | Attach one FIP and open gateway port 8080 | `false` |
| `api_token` | Optional bearer token for gateway auth | `""` |
| `image_name` | Boot image name | `Ubuntu-24.04` |
| `external_network` | Shared external network for router gateway and floating IP | `PublicStatic` |
| `private_cidr` | Private subnet CIDR for the inference instance | `192.168.85.0/24` |
| `instance_name` | Inference instance display name | `creator-ai-worker` |
| `admin_cidr` | Source CIDR allowed for SSH and API access | `0.0.0.0/0` |
| `s3_endpoint` | Quake AI S3-compatible endpoint URL | `https://object.us-east-1.rumble.cloud` |
| `s3_region` | S3 region identifier for the AWS provider | `us-east-1` |

## Cost and sizing

Size `inference_flavor` for the Ollama model you pull and any concurrent whisper transcription. CPU inference suits small models (roughly 7B-8B parameter ceiling on typical flavors). Object Storage charges follow your plan allowance for stored media and Qdrant snapshots.

## When to use this pattern

Wire transcribe-to-summarize-to-index creator workflows (show notes, captions, search) on a single stack without a managed AI SaaS. For batch audio loudness only, see the [audio post-production worker](/resources/iac-templates/audio-worker) template.

## Estimated cost

<PricingCompanion
  components={[
    { kind: "template", slug: "creator-ai-worker", required: true },
  ]}
/>

## Template source

<TemplateSource slug="creator-ai-worker" />

<TemplateResourceMap template="creator-ai-worker" format="opentofu" />

## Outputs

| Output | Description |
| --- | --- |
| `inference_private_ip` | Private IP for bastion/VPN access |
| `inference_floating_ip` | Public IP when `enable_public_api` is true |
| `gateway_url` | Base URL for `/transcribe`, `/summarize`, `/index` |
| `data_bucket` | Object Storage bucket name |
| `floating_ip_count` | `0` by default; `1` when public API is enabled |

## See also

- [Deploy the Creator-AI worker template](/resources/deployments/deploy-creator-ai-worker-template)
- [Self-hosted AI concept](/docs/compute/concepts/self-hosted-ai)
- [Run a local LLM tutorial](/resources/deployments/run-local-llm)
