# Deploy an inference gateway with OpenTofu

Source: https://docs.quake.ai/resources/deployments/deploy-inference-gateway-template
Markdown: https://docs.quake.ai/resources/deployments/deploy-inference-gateway-template.md
> Stand up a LiteLLM model router on one CPU VM with a single OpenAI-compatible HTTPS endpoint using the inference-gateway OpenTofu template.

---

# Deploy an inference gateway with OpenTofu

Stand up [LiteLLM](https://docs.litellm.ai), an open-source model router, on one CPU VM using the [validated OpenTofu template](/docs/platform/validation#how-infrastructure-templates-are-checked) `inference-gateway`. The gateway exposes a single OpenAI-compatible endpoint over HTTPS; model inference runs on the upstream backends you point at.

<PricingCompanion
  components={[
    { kind: "template", slug: "inference-gateway", required: true },
  ]}
/>

<Figure size="md" caption="Inference gateway topology: LiteLLM router with Postgres on a data volume, one floating IP on HTTPS, clients route to upstream backends through one endpoint">

```d2
direction: right

cloud: Quake AI {
  fip: Floating IP\nHTTPS :443
  private: Private network\n10.60.0.0/24 {
    gw: LiteLLM router
    db: Postgres\nkeys + spend log {shape: cylinder}
    vol: Data volume\n/data {shape: cylinder}
    gw -> db: log + keys
    gw -> vol: config + certs
  }
  router: Router\nto PublicStatic
}

backends: Upstream backends\n(elsewhere) {
  primary: Primary model
  backup: Backup model
}

cloud.fip -> cloud.private.gw
cloud.private.gw -> backends.primary: route
cloud.private.gw -> backends.backup: fallback
cloud.router -> cloud.private
```

</Figure>

## Prerequisites

You need:

- A Quake AI account with [application credentials](/docs/tools/generate-app-credentials)
- OpenTofu 1.6.0 or later ([installation guide](https://opentofu.org/docs/intro/install/))
- OpenStack credentials sourced into the shell (`source openrc.sh`). See [the OpenStack CLI guide](/docs/tools/openstack-cli).
- An existing SSH key pair in your project. See [Add an SSH key](/docs/tools/add-ssh-key).
- API keys for at least one upstream model provider. To follow the fallback step, have keys for two OpenAI-compatible backends.
- A copy of the `inference-gateway` template from [the template reference page](/resources/iac-templates/inference-gateway)
- Enough project quota for one `s1a.medium` instance, a 30 GB boot volume, a 20 GB data volume, one router, one private network, and one floating IP

## Step 1: Configure variables and apply

Copy `terraform.tfvars.example` to `terraform.tfvars` and set:

```hcl
key_name = "YOUR_KEY_NAME"
```

Leave `domain` commented out for a self-signed certificate on the floating IP. Defaults for flavor, LiteLLM version, and volume size are documented on the [Inference gateway](/resources/iac-templates/inference-gateway) reference page. Leave `enable_db = true` so the request-logging step has a database to read from.



The default `litellm_version = "main-stable"` tracks the latest stable LiteLLM release. For a deployment you rebuild or hand to a teammate, set a versioned tag such as `litellm_version = "main-v1.77.3-stable"`.



From the template directory, run:

```bash
tofu init
tofu plan
tofu apply
```

Type `yes` when prompted. Provisioning takes a few minutes while cloud-init installs Docker, generates the certificate and master key, and starts the stack.

When the run finishes, record the outputs:

```bash
GATEWAY_URL=$(tofu output -raw gateway_url)
GATEWAY_IP=$(tofu output -raw floating_ip)
```

## Step 2: Retrieve the master key and confirm the gateway answers

Wait for cloud-init to finish, then read the master key the instance generates on first boot:

```bash
ssh -i ~/.ssh/YOUR_KEY -o StrictHostKeyChecking=accept-new ubuntu@"$GATEWAY_IP" 'cloud-init status --wait'
MASTER_KEY=$(ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP" sudo cat /root/gateway-credentials | grep master_key | awk '{print $2}')
```

Copy the first-boot certificate to your workstation:

```bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP" sudo cat /data/certs/gateway.crt > gateway.crt
curl --cacert gateway.crt "$GATEWAY_URL/health/liveliness"
```

A response of `"I'm alive!"` confirms the proxy terminates TLS and answers on 443.

## Step 3: Add upstream provider keys and configure backends

SSH to the instance and edit `/data/config/providers.env`:

```bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP"
sudo nano /data/config/providers.env
```

Set keys for two backends:

```bash
OPENAI_API_KEY=sk-YOUR_PRIMARY_PROVIDER_KEY
UPSTREAM_API_BASE=https://YOUR_BACKUP_ENDPOINT/v1
UPSTREAM_API_KEY=YOUR_BACKUP_PROVIDER_KEY
```

Edit `/data/config/litellm-config.yaml` with two named backends and a fallback:

```yaml
model_list:
  - model_name: chat-primary
    litellm_params:
      model: openai/gpt-4o-mini
      api_key: os.environ/OPENAI_API_KEY
  - model_name: chat-backup
    litellm_params:
      model: openai/YOUR_BACKUP_MODEL
      api_base: os.environ/UPSTREAM_API_BASE
      api_key: os.environ/UPSTREAM_API_KEY

litellm_settings:
  drop_params: true
  cache: true
  cache_params:
    type: local

router_settings:
  routing_strategy: simple-shuffle
  fallbacks: [{"chat-primary": ["chat-backup"]}]

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL
```

Restart the stack and disconnect:

```bash
cd /opt/gateway && sudo docker compose up -d
exit
```

## Step 4: Route requests and confirm fallback

Send a chat completion to each backend through the gateway:

```bash
curl "$GATEWAY_URL/v1/chat/completions" \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "chat-primary", "messages": [{"role": "user", "content": "Reply with one word: hello"}]}'

curl "$GATEWAY_URL/v1/chat/completions" \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "chat-backup", "messages": [{"role": "user", "content": "Reply with one word: hello"}]}'
```

Both calls hit the same base URL with the same key; the `model` field selects which backend serves each request.

To confirm fallback, break the primary key and request `chat-primary` again:

```bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP"
sudo sed -i 's/^OPENAI_API_KEY=.*/OPENAI_API_KEY=sk-invalid-on-purpose/' /data/config/providers.env
cd /opt/gateway && sudo docker compose up -d
exit

curl "$GATEWAY_URL/v1/chat/completions" \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "chat-primary", "messages": [{"role": "user", "content": "Reply with one word: hello"}]}'
```

The response comes back from the backup because the primary key is invalid. Restore the real primary key when you finish.

## Step 5: Confirm request logging in Postgres

With `enable_db = true`, query the spend log from the instance:

```bash
ssh -i ~/.ssh/YOUR_KEY ubuntu@"$GATEWAY_IP"
cd /opt/gateway
sudo docker compose exec db \
  psql -U litellm -d litellm \
  -c 'SELECT request_id, model, total_tokens, spend FROM "LiteLLM_SpendLogs" ORDER BY "startTime" DESC LIMIT 5;'
```

You should see one row per request you sent in the previous steps.

## Next steps

- [Inference gateway template](/resources/iac-templates/inference-gateway)
- [AI inference and RAG](/resources/solutions/ai-inference-rag)
- [Containerized app template](/resources/iac-templates/containerized-app)
- [Redis cache template](/resources/iac-templates/redis-cache) for a shared response cache across replicas

## Clean up

Run `tofu destroy` from the project directory when finished. Type `yes` to confirm. Verify in the Console that the instance and floating IP are gone.
