Skip to content
Deployments

Deploy the Creator-AI inference worker template with OpenTofu

Deployment · Updated Jun 2026

Coming from another cloud?

▸AWS·Bedrock

This Quake AI feature maps to AWS’s Bedrock.

▸Google Cloud·Vertex AI

This Quake AI feature maps to Google Cloud’s Vertex AI.

Deploy the Creator-AI inference worker template with OpenTofu

Stand up a CPU inference VM running Ollama, Qdrant, and a Python gateway for creator workflows using the validated OpenTofu template creator-ai-worker. Keep Ollama models small and expect higher latency than GPU hosts.

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on dedicated vCPU.

Starting template$63.40/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Inference

c2a.large · 2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps

$62.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

c2a.large

2 dedicated vCPU, 4 GiB RAM, 0.5 Gbps

$62.00

Compute + RAM rate basis

2 vCPU + 4 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (80 GiB)

80 GiB at $0.08/GiB/mo

$6.40

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Object storage (usage-based)

Object storage

1 bucket. The first 1 TB is included, then $10.00 per TB each month. You pay for what you store, so this line depends on usage.

$0–$40/mo

Assumes: 2 TB stored is $10/mo; 5 TB stored is $40/mo. Within the included allotment it stays $0.

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Object storage upload and download

No separate charges for uploading or downloading object storage data.

Most object storage providers meter egress and API requests separately from stored capacity.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Configure your estimate

Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.

Starting template

The required baseline, always included.

$63.40/mo

Pick how much you expect to store to fold it into the total.

$0.00/mo
Your configured estimate$63.40/mo

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$17.90/mo

Production on dedicated CPU

The headline estimate above; predictable steady-load performance.

$63.40/mo

Saves $45.50/mo while you build on shared CPU.

Shared flavors carry less RAM (c2a.large (4 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Quake AIObject Storagecaptions + snapshotsInference VMPrivate networkGateway :8080Ollama LLMQdrant vectors summarizeread/write media
Click to zoom
Creator-AI worker topology: private inference VM with gateway API, Ollama, Qdrant, and Object Storage for media

Prerequisites#

You need:

  • A Quake AI account with application credentials
  • OpenTofu 1.6.0 or later (installation guide)
  • curl available for API checks
  • OpenStack credentials sourced into the shell. See the OpenStack CLI guide.
  • A copy of the creator-ai-worker template from the template reference page
  • Enough project quota for one c2a.large instance, one 40 GB boot volume, one private network, and one Object Storage bucket
  • SSH access to the private subnet through a bastion, or set enable_public_api = true for this deployment

Read Self-Hosted AI on Cloud Infrastructure for the platform CPU-only ceiling before you pick an Ollama model tag.

Step 1: Mint credentials and configure variables#

Mint EC2-compatible credentials and export the standard AWS variables:

bash
openstack ec2 credentials create
export AWS_ACCESS_KEY_ID=YOUR_ACCESS_KEY
export AWS_SECRET_ACCESS_KEY=YOUR_SECRET_KEY
export AWS_ENDPOINT_URL_S3=https://object.us-east-1.rumble.cloud

Copy terraform.tfvars.example to terraform.tfvars and set:

HCL
key_name         = "YOUR_KEY_NAME"
data_bucket_name = "YOUR_PROJECT_CREATOR_AI"
s3_access_key    = "YOUR_ACCESS_KEY"
s3_secret_key    = "YOUR_SECRET_KEY"

ollama_model = "llama3.2:1b"

enable_public_api = true
api_token         = "YOUR_GATEWAY_TOKEN"

The template allocates zero floating IPs by default. Setting enable_public_api = true attaches one floating IP and opens port 8080 to admin_cidr.

Step 2: Apply the template#

From the template directory, run:

bash
tofu init
tofu plan
tofu apply

Type yes when prompted. First boot pulls the Ollama model and starts containers. This step can take 10 or more minutes on CPU.

When the run finishes, note gateway_url and data_bucket from the outputs.

Step 3: Verify the gateway#

Call the health route:

bash
curl -sS -H "Authorization: Bearer YOUR_GATEWAY_TOKEN" \
  "$(tofu output -raw gateway_url)/health"

A healthy stack returns JSON with a status field set to ok. If the request times out, confirm cloud-init finished on the inference instance in the Console.

List the data bucket to confirm credentials wired correctly:

bash
aws s3 ls "s3://$(tofu output -raw data_bucket)/" --endpoint-url "$AWS_ENDPOINT_URL_S3"

An empty listing is normal immediately after apply.

Next steps#

Clean up#

Empty the data bucket, then run tofu destroy from the project directory:

bash
aws s3 rm "s3://$(tofu output -raw data_bucket)" --recursive --endpoint-url "$AWS_ENDPOINT_URL_S3"
tofu destroy
Was this page helpful?