# Deploy the Creator-AI inference worker template with OpenTofu

Source: https://docs.quake.ai/resources/deployments/deploy-creator-ai-worker-template
Markdown: https://docs.quake.ai/resources/deployments/deploy-creator-ai-worker-template.md
> Stand up a CPU inference VM running Ollama, Qdrant, and a Python gateway for creator workflows using the creator-ai-worker OpenTofu template.

---

# Deploy the Creator-AI inference worker template with OpenTofu

Stand up a CPU inference VM running Ollama, Qdrant, and a Python gateway for creator workflows using the [validated OpenTofu template](/docs/platform/validation#how-infrastructure-templates-are-checked) `creator-ai-worker`. Keep Ollama models small and expect higher latency than GPU hosts.

<PricingCompanion
  components={[
    { kind: "template", slug: "creator-ai-worker", required: true },
  ]}
/>

<Figure size="md" caption="Creator-AI worker topology: private inference VM with gateway API, Ollama, Qdrant, and Object Storage for media">

```d2
direction: right

cloud: Quake AI {
  bucket: Object Storage\ncaptions + snapshots
  inference: Inference VM {
    gateway: Gateway :8080
    ollama: Ollama LLM
    qdrant: Qdrant vectors
  }
  private: Private network
}

cloud.inference.gateway -> cloud.inference.ollama: summarize
cloud.inference.gateway -> cloud.bucket: read/write media
```

</Figure>

## Prerequisites

You need:

- A Quake AI account with [application credentials](/docs/tools/generate-app-credentials)
- OpenTofu 1.6.0 or later ([installation guide](https://opentofu.org/docs/intro/install/))
- `curl` available for API checks
- OpenStack credentials sourced into the shell. See [the OpenStack CLI guide](/docs/tools/openstack-cli).
- A copy of the `creator-ai-worker` template from [the template reference page](/resources/iac-templates/creator-ai-worker)
- Enough project quota for one `c2a.large` instance, one 40 GB boot volume, one private network, and one Object Storage bucket
- SSH access to the private subnet through a bastion, or set `enable_public_api = true` for this deployment

Read [Self-Hosted AI on Cloud Infrastructure](/docs/compute/concepts/self-hosted-ai) for the platform CPU-only ceiling before you pick an Ollama model tag.

## Step 1: Mint credentials and configure variables

Mint EC2-compatible credentials and export the standard AWS variables:

```bash
openstack ec2 credentials create
export AWS_ACCESS_KEY_ID=YOUR_ACCESS_KEY
export AWS_SECRET_ACCESS_KEY=YOUR_SECRET_KEY
export AWS_ENDPOINT_URL_S3=https://object.us-east-1.rumble.cloud
```

Copy `terraform.tfvars.example` to `terraform.tfvars` and set:

```hcl
key_name         = "YOUR_KEY_NAME"
data_bucket_name = "YOUR_PROJECT_CREATOR_AI"
s3_access_key    = "YOUR_ACCESS_KEY"
s3_secret_key    = "YOUR_SECRET_KEY"

ollama_model = "llama3.2:1b"

enable_public_api = true
api_token         = "YOUR_GATEWAY_TOKEN"
```

The template allocates zero floating IPs by default. Setting `enable_public_api = true` attaches one floating IP and opens port 8080 to `admin_cidr`.

## Step 2: Apply the template

From the template directory, run:

```bash
tofu init
tofu plan
tofu apply
```

Type `yes` when prompted. First boot pulls the Ollama model and starts containers. This step can take 10 or more minutes on CPU.

When the run finishes, note `gateway_url` and `data_bucket` from the outputs.

## Step 3: Verify the gateway

Call the health route:

```bash
curl -sS -H "Authorization: Bearer YOUR_GATEWAY_TOKEN" \
  "$(tofu output -raw gateway_url)/health"
```

A healthy stack returns JSON with a `status` field set to `ok`. If the request times out, confirm cloud-init finished on the inference instance in the Console.

List the data bucket to confirm credentials wired correctly:

```bash
aws s3 ls "s3://$(tofu output -raw data_bucket)/" --endpoint-url "$AWS_ENDPOINT_URL_S3"
```

An empty listing is normal immediately after apply.

## Next steps

- [Creator-AI inference worker template](/resources/iac-templates/creator-ai-worker)
- [Creator businesses](/resources/solutions/creator-businesses)
- [Run a local LLM](/resources/deployments/run-local-llm)
- Disable `enable_public_api` in production and reach the gateway through a bastion or VPN

## Clean up

Empty the data bucket, then run `tofu destroy` from the project directory:

```bash
aws s3 rm "s3://$(tofu output -raw data_bucket)" --recursive --endpoint-url "$AWS_ENDPOINT_URL_S3"
tofu destroy
```
