Inference gateway
Inference gateway
This pattern composes Compute, Network, and Block Storage into a self-hosted model router on infrastructure you control.
What this template does#
Provisions a single instance running LiteLLM, an open-source model router, that exposes one stable OpenAI-compatible endpoint over HTTPS. The gateway routes requests to any model backend you configure: hosted model APIs, a self-hosted vLLM or Ollama endpoint, or any OpenAI-compatible GPU provider. Your clients point a single base URL at the gateway, and the auth, caching, fallback, and request logging run on a VM you own:
- Compute instance that runs the router in Docker, sized for proxy traffic rather than inference
- Private network, subnet, router, port, and security group; a floating IP for public access
- A block volume mounted at
/data, so the routing config, request logs, and the key-and-spend database live on a volume you can grow rather than on the boot disk - cloud-init installs Docker, generates a self-signed TLS certificate and the master key, writes a starter routing config, and brings the stack up on first boot
The model inference itself runs on the upstream backends, so this VM stays small. It is CPU-only, runs in one region, and routes to model and GPU backends that live elsewhere.
No credential ships with this template. The instance generates the master key on first boot and writes it to /root/gateway-credentials (readable only by root); retrieve it over SSH and rotate it after first use. Upstream provider keys are populated by you over SSH into /data/config/providers.env, so they stay out of Terraform state and out of the repo.
With and without the database#
By default the gateway runs an embedded Postgres for virtual API keys, per-key budgets, and spend tracking. Set enable_db = false for a single-process router that authenticates with the master key alone and keeps no request ledger. Use the database when you want to issue scoped keys to teams and track usage per key; the master-key-only mode suits a private gateway with one trusted caller.
Parameters#
| Parameter | Description | Default |
|---|---|---|
key_name | SSH keypair name (must already exist) | No default |
flavor_name | Instance size (4 GiB suits a router fronting several backends) | s1a.medium |
image_name | Operating system image | Ubuntu-24.04 |
app_name | Display name prefix for resources | inference-gateway |
litellm_version | LiteLLM proxy image tag; pin a versioned tag for reproducibility | main-stable |
domain | Public domain for the endpoint; empty uses the floating IP | "" |
volume_size | Block volume size in GiB, mounted at /data | 20 |
enable_db | Run Postgres for virtual keys, budgets, and spend tracking | true |
external_network | External network for floating IP allocation | PublicStatic |
private_cidr | CIDR for the private subnet | 10.60.0.0/24 |
Resource baseline#
The gateway is a router, not an inference host: LiteLLM, the TLS listener, and a small Postgres fit in 4 GiB. The default s1a.medium flavor (2 shared vCPU, 4 GiB RAM) handles steady proxy traffic; raise it if you front many backends or run high request concurrency. The boot disk is 30 GiB; the config, logs, and database live on the separate data volume (volume_size, default 20 GiB). The VM stores no model weights.
Ports and access#
| Port | Purpose |
|---|---|
| 22 | Host SSH for administration, setting provider keys, and retrieving the master key |
| 443 | The one stable gateway endpoint: the OpenAI-compatible API and admin UI over HTTPS |
The proxy serves HTTPS only; there is no plaintext port. On first boot the certificate is self-signed, so a client trusts it by passing /data/certs/gateway.crt as the CA before the first request. For anything exposed to the internet, set domain, point its DNS A record at the floating IP, and replace the first-boot certificate with a CA-issued one. The README in the template directory covers both paths.
How requests flow#
The gateway speaks the OpenAI API. Point any OpenAI-compatible client at gateway_url with the master key (or a virtual key minted from the admin API). The model value in the request matches a model_name from your config, and the gateway forwards the call to the backend that model_name maps to:
curl https://HOST/v1/chat/completions \
--cacert gateway.crt \
-H "Authorization: Bearer $MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "hello"}]}'You configure routes, auth, caching, fallback, and logging in /data/config/litellm-config.yaml on the instance. The starter config documents each point: model_list for the backends a request can map to, general_settings.master_key for auth, litellm_settings.cache for response caching, router_settings for model fallback, and the Postgres-backed spend log for per-request usage.
When to use this pattern#
Run your own model router on a VM you operate, so your apps and agents reach many model backends through one endpoint with one set of keys, instead of wiring each provider into each client. The control plane (the endpoint, the keys, the routing policy, the usage log) stays on infrastructure you own, while the model compute runs on whichever backends you point it at. It pairs with the runtime and data tracks in the IaC library: a containerized app or Next.js app for the agent that calls the gateway, and a Redis cache or self-managed Postgres for application state.
For a deeper treatment of running AI workloads on Quake AI, see AI inference and RAG.
Estimated cost#
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on shared vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Gateway
s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
s1a.medium
4 shared vCPU, 4 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
4 vCPU + 4 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (50 GiB)
50 GiB at $0.08/GiB/mo
Public IP (included)
1 included with the custom package
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Template source#
Show source (7 files)Hide source
data "openstack_images_image_v2" "os" {
name = var.image_name
most_recent = true
}
data "openstack_networking_network_v2" "external" {
name = var.external_network
}
resource "openstack_networking_network_v2" "private" {
name = "${var.app_name}-net"
admin_state_up = true
}
resource "openstack_networking_subnet_v2" "private" {
name = "${var.app_name}-subnet"
network_id = openstack_networking_network_v2.private.id
cidr = var.private_cidr
ip_version = 4
dns_nameservers = ["1.1.1.1", "8.8.8.8"]
}
resource "openstack_networking_router_v2" "main" {
name = "${var.app_name}-router"
external_network_id = data.openstack_networking_network_v2.external.id
}
resource "openstack_networking_router_interface_v2" "private" {
router_id = openstack_networking_router_v2.main.id
subnet_id = openstack_networking_subnet_v2.private.id
}
resource "openstack_networking_secgroup_v2" "gateway" {
name = "${var.app_name}-sg"
description = "Admin SSH and HTTPS for the inference gateway"
}
# Host SSH for administration: populating provider keys and retrieving the
# generated master key.
resource "openstack_networking_secgroup_rule_v2" "ssh" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 22
port_range_max = 22
remote_ip_prefix = "0.0.0.0/0"
security_group_id = openstack_networking_secgroup_v2.gateway.id
}
# HTTPS carries the one stable gateway endpoint: the OpenAI-compatible API and
# the admin UI. The proxy terminates TLS itself; there is no plaintext port.
resource "openstack_networking_secgroup_rule_v2" "https" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 443
port_range_max = 443
remote_ip_prefix = "0.0.0.0/0"
security_group_id = openstack_networking_secgroup_v2.gateway.id
}
resource "openstack_networking_port_v2" "gateway" {
name = "${var.app_name}-port"
network_id = openstack_networking_network_v2.private.id
security_group_ids = [openstack_networking_secgroup_v2.gateway.id]
fixed_ip {
subnet_id = openstack_networking_subnet_v2.private.id
}
depends_on = [openstack_networking_router_interface_v2.private]
}
resource "openstack_blockstorage_volume_v3" "data" {
name = "${var.app_name}-data"
size = var.volume_size
}
resource "openstack_networking_floatingip_v2" "gateway" {
pool = var.external_network
}
resource "openstack_compute_instance_v2" "gateway" {
name = var.app_name
flavor_name = var.flavor_name
key_pair = var.key_name
user_data = templatefile("${path.module}/cloud-init/gateway.yaml.tftpl", {
litellm_version = var.litellm_version
domain = var.domain
enable_db = var.enable_db
floating_ip = openstack_networking_floatingip_v2.gateway.address
})
block_device {
uuid = data.openstack_images_image_v2.os.id
source_type = "image"
destination_type = "volume"
volume_size = 30
boot_index = 0
delete_on_termination = true
}
network {
port = openstack_networking_port_v2.gateway.id
}
}
resource "openstack_compute_volume_attach_v2" "data" {
instance_id = openstack_compute_instance_v2.gateway.id
volume_id = openstack_blockstorage_volume_v3.data.id
}
resource "openstack_networking_floatingip_associate_v2" "gateway" {
floating_ip = openstack_networking_floatingip_v2.gateway.address
port_id = openstack_networking_port_v2.gateway.id
}
variable "key_name" {
description = "SSH keypair name (must already exist in your project)"
type = string
}
variable "flavor_name" {
description = "Instance size. The gateway is a lightweight router: LiteLLM, a reverse-facing TLS listener, and (by default) a small Postgres for key and spend tracking fit comfortably in 4 GiB. The default s1a.medium (2 shared vCPU / 4 GiB) handles steady proxy traffic; raise it if you front many backends or run high request concurrency. Model inference itself runs on the upstream backends, not on this VM."
type = string
default = "s1a.medium"
}
variable "image_name" {
description = "Operating system image. cloud-init targets a Debian-family distribution; Ubuntu 24.04 is the recommended base."
type = string
default = "Ubuntu-24.04"
}
variable "app_name" {
description = "Display name prefix for compute and network resources"
type = string
default = "inference-gateway"
}
variable "litellm_version" {
description = "LiteLLM proxy container image tag from ghcr.io/berriai/litellm. The default main-stable tracks the latest stable release; pin to a versioned tag (for example main-v1.77.3-stable) for reproducible rebuilds. Bump this and re-run cloud-init's bootstrap to upgrade."
type = string
default = "main-stable"
}
variable "domain" {
description = "Public domain for the gateway, used as the TLS hostname and in the endpoint URL (for example gateway.example.com). Leave empty to use the floating IP. The instance always serves HTTPS; with a domain set, point its DNS A record at the floating IP and replace the first-boot self-signed certificate with a CA-issued one (see the README)."
type = string
default = ""
}
variable "volume_size" {
description = "Block volume size in GiB for the gateway config, request logs, and the Postgres database. The volume is mounted at /data so gateway state lives on a resizable volume rather than the boot disk. The router stores no model weights; 20 GiB is ample for config and logs."
type = number
default = 20
}
variable "enable_db" {
description = "Run an embedded Postgres alongside the proxy for virtual API keys, per-key budgets, and spend tracking. Set false for a stateless router that authenticates with the master key only and keeps no request ledger."
type = bool
default = true
}
variable "external_network" {
description = "Shared external network for router gateway and floating IPs; defaults to PublicStatic (persisted FIP / production pattern). Override with PublicEphemeral for ephemeral demos."
type = string
default = "PublicStatic"
}
variable "private_cidr" {
description = "CIDR for the private tenant network the instance lives in"
type = string
default = "10.60.0.0/24"
}
output "instance_id" {
description = "ID of the compute instance running the gateway"
value = openstack_compute_instance_v2.gateway.id
}
output "floating_ip" {
description = "Public floating IP address of the gateway"
value = openstack_networking_floatingip_v2.gateway.address
}
output "private_ip" {
description = "Private IP address of the instance"
value = openstack_compute_instance_v2.gateway.access_ip_v4
}
output "gateway_url" {
description = "The one stable gateway endpoint. Clients point their OpenAI-compatible base URL here. Uses the domain when set, otherwise the floating IP. The first-boot certificate is self-signed; replace it with a CA-issued certificate before relying on the domain (see the README)."
value = var.domain != "" ? "https://${var.domain}" : "https://${openstack_networking_floatingip_v2.gateway.address}"
}
terraform {
required_version = ">= 1.6.0"
required_providers {
openstack = {
source = "terraform-provider-openstack/openstack"
version = "~> 2.0"
}
}
}
provider "openstack" {}
# Required: SSH keypair must already exist in your project
key_name = "YOUR_KEY_NAME"
# Recommended: set a domain so the endpoint URL and TLS certificate use a stable
# name. Point its DNS A record at the floating IP and replace the first-boot
# self-signed certificate with a CA-issued one (see the README).
# domain = "gateway.example.com"
# litellm_version = "main-stable"
# flavor_name = "s1a.medium"
# image_name = "Ubuntu-24.04"
# app_name = "inference-gateway"
# volume_size = 20
# enable_db = true
# external_network = "PublicStatic"
# private_cidr = "10.60.0.0/24"
# Upstream provider API keys are NOT set here: they would land in Terraform
# state. After apply, SSH to the instance and populate /data/config/providers.env,
# then restart the stack. See the README, "Provider keys and the master key".
#cloud-config
package_update: true
packages:
- ca-certificates
- curl
- openssl
write_files:
# First-boot bootstrap. Generates a self-signed TLS certificate and the
# LiteLLM master key on the instance (and a Postgres password when the
# database is enabled): none of them is shipped in this template. Writes a
# starter LiteLLM config, an empty provider-keys file, and a docker compose
# file, then brings the stack up. The master key is written to
# /root/gateway-credentials (root-only); retrieve it over SSH.
- path: /opt/gateway-bootstrap.sh
permissions: "0755"
content: |
#!/usr/bin/env bash
set -euo pipefail
CRED_FILE=/root/gateway-credentials
STACK_DIR=/opt/gateway
CONFIG_DIR=/data/config
CERT_DIR=/data/certs
GW_HOST="${domain != "" ? domain : floating_ip}"
# Idempotent across reboots: if already installed, just bring the stack up.
if [ -f "$CRED_FILE" ]; then
cd "$STACK_DIR"
docker compose up -d
exit 0
fi
mkdir -p "$CERT_DIR" "$CONFIG_DIR" "$STACK_DIR" /data/postgres
# First-boot self-signed certificate. The proxy serves HTTPS only, so
# clients need TLS from the first request. This cert lets them connect
# after trusting it; replace it with a CA-issued certificate for anything
# beyond a private network or a trial (see the README).
openssl req -x509 -newkey rsa:4096 -nodes \
-keyout "$CERT_DIR/gateway.key" \
-out "$CERT_DIR/gateway.crt" \
-days 825 \
-subj "/CN=$GW_HOST" \
-addext "subjectAltName=${domain != "" ? "DNS:${domain}" : "IP:${floating_ip}"}"
chmod 600 "$CERT_DIR/gateway.key"
# Generate the master key (clients authenticate to the gateway with it) and
# the database password on the instance. Neither is shipped in the template.
MASTER_KEY="sk-$(openssl rand -hex 24)"
DB_PASSWORD="$(openssl rand -base64 24 | tr -dc 'A-Za-z0-9' | head -c 24)"
# Starter routing config. model_name is what clients request; the
# litellm_params.model is the provider/model it maps to. Provider keys come
# from providers.env via os.environ, so no secret lives in this file.
cat > "$CONFIG_DIR/litellm-config.yaml" <<'LITELLM_CONFIG'
model_list:
# Duplicate and edit one entry per upstream backend you route to (OpenAI,
# Anthropic, a self-hosted vLLM/Ollama endpoint, or any OpenAI-compatible
# GPU provider). The model inference runs on those backends; this VM only
# routes requests to them.
- model_name: gpt-4o-mini
litellm_params:
model: openai/gpt-4o-mini
api_key: os.environ/OPENAI_API_KEY
- model_name: my-openai-compatible-model
litellm_params:
model: openai/your-model-name
api_base: os.environ/UPSTREAM_API_BASE
api_key: os.environ/UPSTREAM_API_KEY
litellm_settings:
drop_params: true
# Response caching cuts cost and latency on repeated prompts. The local
# in-memory cache resets on restart; point cache_params at a Redis backend
# for a cache that survives restarts and is shared across replicas.
cache: true
cache_params:
type: local
router_settings:
# Fallback: when a model errors or rate-limits, route to a backup. Map a
# primary model_name to an ordered list of backups, for example:
# fallbacks: [{"gpt-4o-mini": ["my-openai-compatible-model"]}]
routing_strategy: simple-shuffle
general_settings:
# Auth: every client request must carry this key as a bearer token. Mint
# scoped virtual keys from the admin API when the database is enabled.
master_key: os.environ/LITELLM_MASTER_KEY
%{ if enable_db }
# Logging and spend tracking persist to this database.
database_url: os.environ/DATABASE_URL
%{ endif }
LITELLM_CONFIG
# Provider keys live here, not in the template or in Terraform state. Fill
# in the keys your config references, then: cd /opt/gateway && docker compose up -d
cat > "$CONFIG_DIR/providers.env" <<'PROVIDERS_ENV'
# Upstream provider API keys. Uncomment and set the ones your
# litellm-config.yaml references via os.environ, then restart the stack.
# OPENAI_API_KEY=
# ANTHROPIC_API_KEY=
# For a generic OpenAI-compatible backend (self-hosted vLLM, a GPU provider):
# UPSTREAM_API_BASE=
# UPSTREAM_API_KEY=
PROVIDERS_ENV
chmod 600 "$CONFIG_DIR/providers.env"
# Docker compose for the gateway. The proxy terminates TLS itself and
# publishes the one stable endpoint on 443; model inference happens on the
# upstream backends, never on this VM.
cat > "$STACK_DIR/docker-compose.yml" <<COMPOSE
services:
litellm:
image: ghcr.io/berriai/litellm:${litellm_version}
restart: unless-stopped
ports:
- "443:4000"
env_file:
- /data/config/providers.env
environment:
- LITELLM_MASTER_KEY=$MASTER_KEY
%{ if enable_db }
- DATABASE_URL=postgresql://litellm:$DB_PASSWORD@db:5432/litellm
%{ endif }
volumes:
- /data/config/litellm-config.yaml:/app/config.yaml:ro
- /data/certs:/data/certs:ro
command: ["--config", "/app/config.yaml", "--port", "4000", "--ssl_certfile_path", "/data/certs/gateway.crt", "--ssl_keyfile_path", "/data/certs/gateway.key"]
%{ if enable_db }
depends_on:
- db
db:
image: postgres:16
restart: unless-stopped
environment:
- POSTGRES_USER=litellm
- POSTGRES_PASSWORD=$DB_PASSWORD
- POSTGRES_DB=litellm
volumes:
- /data/postgres:/var/lib/postgresql/data
%{ endif }
COMPOSE
cd "$STACK_DIR"
docker compose up -d
umask 077
cat > "$CRED_FILE" <<CRED
Inference gateway master key (generated on first boot)
endpoint: ${domain != "" ? "https://${domain}" : "https://${floating_ip}"}
master_key: $MASTER_KEY
Clients send this as a bearer token. Mint scoped virtual keys from the admin
API instead of sharing this one, then rotate it and delete this file.
CRED
chmod 600 "$CRED_FILE"
runcmd:
- |
set -e
# The data volume attaches as /dev/sdb on this platform (not /dev/vdb).
# Mount it at /data before the stack runs, so config, request logs, and the
# database live on the resizable volume rather than the boot disk.
DEV=/dev/sdb
for i in $(seq 1 30); do [ -b "$DEV" ] && break; sleep 5; done
if ! blkid "$DEV" >/dev/null 2>&1; then mkfs.ext4 -F -L gatewaydata "$DEV"; fi
mkdir -p /data
mount "$DEV" /data
grep -q "$DEV" /etc/fstab || echo "$DEV /data ext4 defaults,nofail 0 2" >> /etc/fstab
# Install Docker Engine (provides docker + the compose plugin).
curl -fsSL https://get.docker.com | sh
systemctl enable --now docker
/opt/gateway-bootstrap.sh
# Inference gateway
Single compute instance running [LiteLLM](https://docs.litellm.ai), an open-source model router, on infrastructure you control. The gateway exposes one stable OpenAI-compatible endpoint over HTTPS and routes requests to any model backend you configure: hosted model APIs, a self-hosted vLLM or Ollama endpoint, or any OpenAI-compatible GPU provider. Clients point a single base URL at the gateway; auth, caching, fallback, and request logging live on the VM you own, and the heavy model inference runs on the upstream backends.
**Network class:** production — `external_network` defaults to `PublicStatic` for persisted floating IPs and multi-tier stacks; override with `PublicEphemeral` for ephemeral demos.
The instance provisions a private network, a floating IP, and a block volume mounted at `/data` so the config, request logs, and (by default) the key-and-spend database live on a resizable volume rather than the boot disk. cloud-init installs Docker, generates a self-signed TLS certificate and the master key, writes a starter routing config, and brings the stack up, all on first boot.
## Where this fits
This is the control-plane side of an AI workload: the always-on CPU layer that holds provider credentials, routing policy, and a single endpoint, while the model compute lives on backends that can be anywhere. It pairs with the data and runtime tracks in the IaC library: a [Redis cache](/resources/iac-templates/redis-cache) or [self-managed Postgres](/resources/iac-templates/self-managed-postgres) for app state, and a [containerized app](/resources/iac-templates/containerized-app) or [Next.js app](/resources/iac-templates/nextjs-app) for the agent or application that calls the gateway.
For a lighter, single-process router with no admin database, set `enable_db = false`; the gateway then authenticates with the master key only and keeps no request ledger.
## Prerequisites
- OpenTofu >= 1.6.0 or Terraform >= 1.6.0
- Quake AI account with OpenStack credentials
- An existing SSH keypair in your project (the value of `key_name` must match that keypair)
- API keys for the upstream model providers you intend to route to
## Resource baseline
The gateway is a lightweight router, not an inference host: LiteLLM, the TLS listener, and a small Postgres fit in 4 GiB. The default `s1a.medium` flavor (2 shared vCPU, 4 GiB RAM) handles steady proxy traffic; raise it if you front many backends or run high request concurrency. The boot disk is 30 GiB; config, logs, and the database live on the separate data volume (`volume_size`, default 20 GiB). No model weights are stored on this VM.
## Usage
1. Clone or copy this template directory
2. Copy `terraform.tfvars.example` to `terraform.tfvars` and set `key_name` (and `domain` if you have one)
3. Source your OpenStack credentials: `source openrc.sh`
4. Initialize: `tofu init`
5. Preview: `tofu plan`
6. Apply: `tofu apply`
cloud-init takes a few minutes on first boot to install Docker, pull the proxy image, and start the stack. The gateway then answers on `gateway_url` from the outputs, but it cannot reach any backend until you add provider keys (next section).
## Provider keys and the master key
No credential ships with this template, and provider keys never pass through Terraform state.
- **The master key** (how clients authenticate to the gateway) is generated on first boot and written to `/root/gateway-credentials` (mode 600). Retrieve it over SSH: `ssh user@<floating_ip> sudo cat /root/gateway-credentials`. Mint scoped virtual keys from the admin API instead of sharing the master key, then rotate it.
- **Upstream provider keys** are populated by you after apply, so they stay out of state and out of the repo:
```bash
ssh user@<floating_ip>
sudo nano /data/config/providers.env # set OPENAI_API_KEY, UPSTREAM_API_KEY, etc.
cd /opt/gateway && sudo docker compose up -d # restart to pick up the keys
```
The starter `/data/config/litellm-config.yaml` references those keys as `os.environ/OPENAI_API_KEY` and similar, so the secret stays in the env file, never in the config.
## Configuring routes, auth, caching, fallback, and logging
Edit `/data/config/litellm-config.yaml` on the instance and restart the stack to apply. The starter config documents each point:
- **Routes** (`model_list`): one entry per upstream model. `model_name` is what clients request; `litellm_params.model` is the provider/model it maps to. Add an `api_base` for self-hosted or third-party OpenAI-compatible endpoints.
- **Auth** (`general_settings.master_key`): every request carries the key as a bearer token. With the database enabled, mint per-team virtual keys with their own budgets from the admin API.
- **Caching** (`litellm_settings.cache`): a local in-memory cache is on by default; point `cache_params` at a Redis backend for a cache that survives restarts.
- **Fallback** (`router_settings`): map a primary `model_name` to an ordered list of backups so a model error or rate-limit routes to the next backend.
- **Logging**: the proxy logs to stdout (`docker compose logs litellm`); with the database enabled, per-request spend and usage persist to Postgres and surface in the admin UI.
## Ports and access
| Port | Purpose | Open to |
| --- | --- | --- |
| 22 | Host SSH for administration, setting provider keys, and retrieving the master key | `0.0.0.0/0` |
| 443 | The one stable gateway endpoint: the OpenAI-compatible API and admin UI over HTTPS | `0.0.0.0/0` |
The proxy serves HTTPS only; there is no plaintext port.
## Calling the gateway
The gateway speaks the OpenAI API. Point any OpenAI-compatible client at `gateway_url` with the master key (or a virtual key). Because the first-boot certificate is self-signed, a client either trusts the certificate or, for a quick test, skips verification:
```bash
curl https://HOST/v1/chat/completions \
--cacert gateway.crt \
-H "Authorization: Bearer $MASTER_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "hello"}]}'
```
`HOST` is your domain or the floating IP; `gateway.crt` is the certificate copied from `/data/certs/gateway.crt`. The `model` value matches a `model_name` from your config.
## TLS for production
The first-boot certificate is self-signed, which is fine for a private network or a quick trial. For anything exposed to the internet:
- Set `domain`, point its DNS A record at `floating_ip`, and obtain a CA-issued certificate (for example with `certbot certonly`).
- Replace `/data/certs/gateway.crt` and `/data/certs/gateway.key` with the issued certificate and key, then `cd /opt/gateway && docker compose up -d` to apply.
With a CA-issued certificate, clients trust the gateway without the `--cacert` step above.
## Variables
| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `key_name` | string | yes | n/a | SSH keypair name (must already exist in your project) |
| `flavor_name` | string | no | `s1a.medium` | Instance size (4 GiB suits a router fronting several backends) |
| `image_name` | string | no | `Ubuntu-24.04` | Operating system image |
| `app_name` | string | no | `inference-gateway` | Display name prefix for resources |
| `litellm_version` | string | no | `main-stable` | LiteLLM proxy image tag; pin a versioned tag for reproducibility |
| `domain` | string | no | `""` | Public domain for the endpoint; empty uses the floating IP |
| `volume_size` | number | no | `20` | Block volume size in GiB, mounted at `/data` |
| `enable_db` | bool | no | `true` | Run Postgres for virtual keys, budgets, and spend tracking |
| `external_network` | string | no | `PublicStatic` | Persisted FIP / production default; override with `PublicEphemeral` for demos |
| `private_cidr` | string | no | `10.60.0.0/24` | CIDR for the private subnet |
## Outputs
| Name | Description |
| --- | --- |
| `floating_ip` | Public floating IP assigned to the instance |
| `private_ip` | Private IP address of the instance |
| `gateway_url` | The one stable gateway endpoint (domain when set, otherwise the floating IP) |
| `instance_id` | Compute instance ID |
## Scope
This is a single-VM gateway you operate, not a managed service. It is CPU-only and runs in one region; it routes to model and GPU backends that live elsewhere and makes no first-party GPU assumption. For high availability, run LiteLLM against external Postgres and Redis and place multiple instances behind a load balancer; this template provisions one node with its database and cache on the attached volume.
## Documentation
See also: [Redis cache template](/resources/iac-templates/redis-cache), [Containerized app template](/resources/iac-templates/containerized-app)
Resources, parameters, and variables
key_namerequiredflavor_name="s1a.medium"image_name="Ubuntu-24.04"app_name="inference-gateway"litellm_version="main-stable"domain=""volume_size=20enable_db=trueexternal_network="PublicStatic"private_cidr="10.60.0.0/24"
Customize this pattern#
- Customize a template's image and flavor
- Add a block volume to a template
- Parameterize a template with a tfvars file
See also#
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.
Last validated: 29.06.2026
See Also
Terraform and OpenTofu on Quake AI
Prerequisite
Networks
Prerequisite
Authoring IaC templates for Quake AI
Shares: Volumes, Security Groups
Deploy an API gateway with the api-gateway template
Shares: Volumes, Security Groups
Deploy a regional edge cache with the edge-cache template
Shares: Volumes, Security Groups