Skip to content

Deploy an inference gateway with OpenTofu

Deployment

Coming from another cloud?

▸AWS·Bedrock

This Quake AI feature maps to AWS’s Bedrock.

▸Google Cloud·Vertex AI

This Quake AI feature maps to Google Cloud’s Vertex AI.

Deploy an inference gateway with OpenTofu

Stand up LiteLLM, an open-source model router, on one CPU VM using the validated OpenTofu template inference-gateway. The gateway exposes a single OpenAI-compatible endpoint over HTTPS; model inference runs on the upstream backends you point at.

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on shared vCPU.

Starting template$32.00/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Gateway

s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

Compute + RAM rate basis

4 vCPU + 4 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (50 GiB)

50 GiB at $0.08/GiB/mo

$4.00

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Pricing data last validated: . For current rates, check quake.ai/pricing.

Quake AIUpstream backends(elsewhere)Floating IPHTTPS :443Private network10.60.0.0/24Routerto PublicStaticPrimary modelBackup modelLiteLLM routerPostgreskeys + spend logData volume/data log + keysconfig + certsroutefallback
Click to zoom
Inference gateway topology: LiteLLM router with Postgres on a data volume, one floating IP on HTTPS, clients route to upstream backends through one endpoint

Prerequisites#

You need:

  • A Quake AI account with application credentials
  • OpenTofu 1.6.0 or later (installation guide)
  • OpenStack credentials sourced into the shell (source openrc.sh). See the OpenStack CLI guide.
  • An existing SSH key pair in your project. See Add an SSH key.
  • API keys for at least one upstream model provider. To follow the fallback step, have keys for two OpenAI-compatible backends.
  • A copy of the inference-gateway template from the template reference page
  • Enough project quota for one s1a.medium instance, a 30 GB boot volume, a 20 GB data volume, one router, one private network, and one floating IP

Step 1: Configure variables and apply#

Copy terraform.tfvars.example to terraform.tfvars and set:

HCL
key_name = "YOUR_KEY_NAME"

Leave domain commented out for a self-signed certificate on the floating IP. Defaults for flavor, LiteLLM version, and volume size are documented on the Inference gateway reference page. Leave enable_db = true so the request-logging step has a database to read from.

From the template directory, run:

bash
tofu init
tofu plan
tofu apply

Type yes when prompted. Provisioning takes a few minutes while cloud-init installs Docker, generates the certificate and master key, and starts the stack.

When the run finishes, record the outputs:

bash
GATEWAY_URL=$(tofu output -raw gateway_url)
GATEWAY_IP=$(tofu output -raw floating_ip)

Step 2: Retrieve the master key and confirm the gateway answers#

Wait for cloud-init to finish, then read the master key the instance generates on first boot:

bash
ssh -i ~/.ssh/YOUR_KEY -o StrictHostKeyChecking=accept-new ubuntu@"$GATEWAY_IP" 'cloud-init status --wait'
MASTER_KEY=$(ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP" sudo cat /root/gateway-credentials | grep master_key | awk '{print $2}')

Copy the first-boot certificate to your workstation:

bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP" sudo cat /data/certs/gateway.crt > gateway.crt
curl --cacert gateway.crt "$GATEWAY_URL/health/liveliness"

A response of "I'm alive!" confirms the proxy terminates TLS and answers on 443.

Step 3: Add upstream provider keys and configure backends#

SSH to the instance and edit /data/config/providers.env:

bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP"
sudo nano /data/config/providers.env

Set keys for two backends:

bash
OPENAI_API_KEY=sk-YOUR_PRIMARY_PROVIDER_KEY
UPSTREAM_API_BASE=https://YOUR_BACKUP_ENDPOINT/v1
UPSTREAM_API_KEY=YOUR_BACKUP_PROVIDER_KEY

Edit /data/config/litellm-config.yaml with two named backends and a fallback:

YAML
model_list:
  - model_name: chat-primary
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: chat-backup
    litellm_params:
      model: openai/YOUR_BACKUP_MODEL
      api_base: os.environ/UPSTREAM_API_BASE
      api_key: os.environ/UPSTREAM_API_KEY

litellm_settings:
  drop_params: true
  cache: true
  cache_params:
    type: local

router_settings:
  routing_strategy: simple-shuffle
  fallbacks: [{"chat-primary": ["chat-backup"]}]

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL

Restart the stack and disconnect:

bash
cd /opt/gateway && sudo docker compose up -d
exit

Step 4: Route requests and confirm fallback#

Send a chat completion to each backend through the gateway:

bash
curl "$GATEWAY_URL/v1/chat/completions" \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "chat-primary", "messages": [{"role": "user", "content": "Reply with one word: hello"}]}'

curl "$GATEWAY_URL/v1/chat/completions" \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "chat-backup", "messages": [{"role": "user", "content": "Reply with one word: hello"}]}'

Both calls hit the same base URL with the same key; the model field selects which backend serves each request.

To confirm fallback, break the primary key and request chat-primary again:

bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP"
sudo sed -i 's/^OPENAI_API_KEY=.*/OPENAI_API_KEY=sk-invalid-on-purpose/' /data/config/providers.env
cd /opt/gateway && sudo docker compose up -d
exit

curl "$GATEWAY_URL/v1/chat/completions" \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "chat-primary", "messages": [{"role": "user", "content": "Reply with one word: hello"}]}'

The response comes back from the backup because the primary key is invalid. Restore the real primary key when you finish.

Step 5: Confirm request logging in Postgres#

With enable_db = true, query the spend log from the instance:

bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP"
cd /opt/gateway
sudo docker compose exec db \
  psql -U litellm -d litellm \
  -c 'SELECT request_id, model, total_tokens, spend FROM "LiteLLM_SpendLogs" ORDER BY "startTime" DESC LIMIT 5;'

You should see one row per request you sent in the previous steps.

Next steps#

Clean up#

Run tofu destroy from the project directory when finished. Type yes to confirm. Verify in the Console that the instance and floating IP are gone.

Was this page helpful?