Skip to content

Inference gateway

Template · Updated Jun 2026
Validated Jun 2026

Inference gateway

This pattern composes Compute, Network, and Block Storage into a self-hosted model router on infrastructure you control.

What this template does#

Provisions a single instance running LiteLLM, an open-source model router, that exposes one stable OpenAI-compatible endpoint over HTTPS. The gateway routes requests to any model backend you configure: hosted model APIs, a self-hosted vLLM or Ollama endpoint, or any OpenAI-compatible GPU provider. Your clients point a single base URL at the gateway, and the auth, caching, fallback, and request logging run on a VM you own:

  • Compute instance that runs the router in Docker, sized for proxy traffic rather than inference
  • Private network, subnet, router, port, and security group; a floating IP for public access
  • A block volume mounted at /data, so the routing config, request logs, and the key-and-spend database live on a volume you can grow rather than on the boot disk
  • cloud-init installs Docker, generates a self-signed TLS certificate and the master key, writes a starter routing config, and brings the stack up on first boot

The model inference itself runs on the upstream backends, so this VM stays small. It is CPU-only, runs in one region, and routes to model and GPU backends that live elsewhere.

No credential ships with this template. The instance generates the master key on first boot and writes it to /root/gateway-credentials (readable only by root); retrieve it over SSH and rotate it after first use. Upstream provider keys are populated by you over SSH into /data/config/providers.env, so they stay out of Terraform state and out of the repo.

With and without the database#

By default the gateway runs an embedded Postgres for virtual API keys, per-key budgets, and spend tracking. Set enable_db = false for a single-process router that authenticates with the master key alone and keeps no request ledger. Use the database when you want to issue scoped keys to teams and track usage per key; the master-key-only mode suits a private gateway with one trusted caller.

Parameters#

ParameterDescriptionDefault
key_nameSSH keypair name (must already exist)No default
flavor_nameInstance size (4 GiB suits a router fronting several backends)s1a.medium
image_nameOperating system imageUbuntu-24.04
app_nameDisplay name prefix for resourcesinference-gateway
litellm_versionLiteLLM proxy image tag; pin a versioned tag for reproducibilitymain-stable
domainPublic domain for the endpoint; empty uses the floating IP""
volume_sizeBlock volume size in GiB, mounted at /data20
enable_dbRun Postgres for virtual keys, budgets, and spend trackingtrue
external_networkExternal network for floating IP allocationPublicStatic
private_cidrCIDR for the private subnet10.60.0.0/24

Resource baseline#

The gateway is a router, not an inference host: LiteLLM, the TLS listener, and a small Postgres fit in 4 GiB. The default s1a.medium flavor (2 shared vCPU, 4 GiB RAM) handles steady proxy traffic; raise it if you front many backends or run high request concurrency. The boot disk is 30 GiB; the config, logs, and database live on the separate data volume (volume_size, default 20 GiB). The VM stores no model weights.

Ports and access#

PortPurpose
22Host SSH for administration, setting provider keys, and retrieving the master key
443The one stable gateway endpoint: the OpenAI-compatible API and admin UI over HTTPS

The proxy serves HTTPS only; there is no plaintext port. On first boot the certificate is self-signed, so a client trusts it by passing /data/certs/gateway.crt as the CA before the first request. For anything exposed to the internet, set domain, point its DNS A record at the floating IP, and replace the first-boot certificate with a CA-issued one. The README in the template directory covers both paths.

How requests flow#

The gateway speaks the OpenAI API. Point any OpenAI-compatible client at gateway_url with the master key (or a virtual key minted from the admin API). The model value in the request matches a model_name from your config, and the gateway forwards the call to the backend that model_name maps to:

bash
curl https://HOST/v1/chat/completions \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "hello"}]}'

You configure routes, auth, caching, fallback, and logging in /data/config/litellm-config.yaml on the instance. The starter config documents each point: model_list for the backends a request can map to, general_settings.master_key for auth, litellm_settings.cache for response caching, router_settings for model fallback, and the Postgres-backed spend log for per-request usage.

When to use this pattern#

Run your own model router on a VM you operate, so your apps and agents reach many model backends through one endpoint with one set of keys, instead of wiring each provider into each client. The control plane (the endpoint, the keys, the routing policy, the usage log) stays on infrastructure you own, while the model compute runs on whichever backends you point it at. It pairs with the runtime and data tracks in the IaC library: a containerized app or Next.js app for the agent that calls the gateway, and a Redis cache or self-managed Postgres for application state.

For a deeper treatment of running AI workloads on Quake AI, see AI inference and RAG.

Estimated cost#

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on shared vCPU.

Starting template$32.00/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Gateway

s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

Compute + RAM rate basis

4 vCPU + 4 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (50 GiB)

50 GiB at $0.08/GiB/mo

$4.00

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Pricing data last validated: . For current rates, check quake.ai/pricing.

Template source#

7 files. Download the zip or expand to copy any file.Download inference-gateway.zip
Show source (7 files)
main.tfHCL
data "openstack_images_image_v2" "os" {
  name        = var.image_name
  most_recent = true
}

data "openstack_networking_network_v2" "external" {
  name = var.external_network
}

resource "openstack_networking_network_v2" "private" {
  name           = "${var.app_name}-net"
  admin_state_up = true
}

resource "openstack_networking_subnet_v2" "private" {
  name            = "${var.app_name}-subnet"
  network_id      = openstack_networking_network_v2.private.id
  cidr            = var.private_cidr
  ip_version      = 4
  dns_nameservers = ["1.1.1.1", "8.8.8.8"]
}

resource "openstack_networking_router_v2" "main" {
  name                = "${var.app_name}-router"
  external_network_id = data.openstack_networking_network_v2.external.id
}

resource "openstack_networking_router_interface_v2" "private" {
  router_id = openstack_networking_router_v2.main.id
  subnet_id = openstack_networking_subnet_v2.private.id
}

resource "openstack_networking_secgroup_v2" "gateway" {
  name        = "${var.app_name}-sg"
  description = "Admin SSH and HTTPS for the inference gateway"
}

# Host SSH for administration: populating provider keys and retrieving the
# generated master key.
resource "openstack_networking_secgroup_rule_v2" "ssh" {
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 22
  port_range_max    = 22
  remote_ip_prefix  = "0.0.0.0/0"
  security_group_id = openstack_networking_secgroup_v2.gateway.id
}

# HTTPS carries the one stable gateway endpoint: the OpenAI-compatible API and
# the admin UI. The proxy terminates TLS itself; there is no plaintext port.
resource "openstack_networking_secgroup_rule_v2" "https" {
  direction         = "ingress"
  ethertype         = "IPv4"
  protocol          = "tcp"
  port_range_min    = 443
  port_range_max    = 443
  remote_ip_prefix  = "0.0.0.0/0"
  security_group_id = openstack_networking_secgroup_v2.gateway.id
}

resource "openstack_networking_port_v2" "gateway" {
  name               = "${var.app_name}-port"
  network_id         = openstack_networking_network_v2.private.id
  security_group_ids = [openstack_networking_secgroup_v2.gateway.id]

  fixed_ip {
    subnet_id = openstack_networking_subnet_v2.private.id
  }

  depends_on = [openstack_networking_router_interface_v2.private]
}

resource "openstack_blockstorage_volume_v3" "data" {
  name = "${var.app_name}-data"
  size = var.volume_size
}

resource "openstack_networking_floatingip_v2" "gateway" {
  pool = var.external_network
}

resource "openstack_compute_instance_v2" "gateway" {
  name        = var.app_name
  flavor_name = var.flavor_name
  key_pair    = var.key_name

  user_data = templatefile("${path.module}/cloud-init/gateway.yaml.tftpl", {
    litellm_version = var.litellm_version
    domain          = var.domain
    enable_db       = var.enable_db
    floating_ip     = openstack_networking_floatingip_v2.gateway.address
  })

  block_device {
    uuid                  = data.openstack_images_image_v2.os.id
    source_type           = "image"
    destination_type      = "volume"
    volume_size           = 30
    boot_index            = 0
    delete_on_termination = true
  }

  network {
    port = openstack_networking_port_v2.gateway.id
  }
}

resource "openstack_compute_volume_attach_v2" "data" {
  instance_id = openstack_compute_instance_v2.gateway.id
  volume_id   = openstack_blockstorage_volume_v3.data.id
}

resource "openstack_networking_floatingip_associate_v2" "gateway" {
  floating_ip = openstack_networking_floatingip_v2.gateway.address
  port_id     = openstack_networking_port_v2.gateway.id
}
variables.tfHCL
variable "key_name" {
  description = "SSH keypair name (must already exist in your project)"
  type        = string
}

variable "flavor_name" {
  description = "Instance size. The gateway is a lightweight router: LiteLLM, a reverse-facing TLS listener, and (by default) a small Postgres for key and spend tracking fit comfortably in 4 GiB. The default s1a.medium (2 shared vCPU / 4 GiB) handles steady proxy traffic; raise it if you front many backends or run high request concurrency. Model inference itself runs on the upstream backends, not on this VM."
  type        = string
  default     = "s1a.medium"
}

variable "image_name" {
  description = "Operating system image. cloud-init targets a Debian-family distribution; Ubuntu 24.04 is the recommended base."
  type        = string
  default     = "Ubuntu-24.04"
}

variable "app_name" {
  description = "Display name prefix for compute and network resources"
  type        = string
  default     = "inference-gateway"
}

variable "litellm_version" {
  description = "LiteLLM proxy container image tag from ghcr.io/berriai/litellm. The default main-stable tracks the latest stable release; pin to a versioned tag (for example main-v1.77.3-stable) for reproducible rebuilds. Bump this and re-run cloud-init's bootstrap to upgrade."
  type        = string
  default     = "main-stable"
}

variable "domain" {
  description = "Public domain for the gateway, used as the TLS hostname and in the endpoint URL (for example gateway.example.com). Leave empty to use the floating IP. The instance always serves HTTPS; with a domain set, point its DNS A record at the floating IP and replace the first-boot self-signed certificate with a CA-issued one (see the README)."
  type        = string
  default     = ""
}

variable "volume_size" {
  description = "Block volume size in GiB for the gateway config, request logs, and the Postgres database. The volume is mounted at /data so gateway state lives on a resizable volume rather than the boot disk. The router stores no model weights; 20 GiB is ample for config and logs."
  type        = number
  default     = 20
}

variable "enable_db" {
  description = "Run an embedded Postgres alongside the proxy for virtual API keys, per-key budgets, and spend tracking. Set false for a stateless router that authenticates with the master key only and keeps no request ledger."
  type        = bool
  default     = true
}

variable "external_network" {
  description = "Shared external network for router gateway and floating IPs; defaults to PublicStatic (persisted FIP / production pattern). Override with PublicEphemeral for ephemeral demos."
  type        = string
  default     = "PublicStatic"
}

variable "private_cidr" {
  description = "CIDR for the private tenant network the instance lives in"
  type        = string
  default     = "10.60.0.0/24"
}
outputs.tfHCL
output "instance_id" {
  description = "ID of the compute instance running the gateway"
  value       = openstack_compute_instance_v2.gateway.id
}

output "floating_ip" {
  description = "Public floating IP address of the gateway"
  value       = openstack_networking_floatingip_v2.gateway.address
}

output "private_ip" {
  description = "Private IP address of the instance"
  value       = openstack_compute_instance_v2.gateway.access_ip_v4
}

output "gateway_url" {
  description = "The one stable gateway endpoint. Clients point their OpenAI-compatible base URL here. Uses the domain when set, otherwise the floating IP. The first-boot certificate is self-signed; replace it with a CA-issued certificate before relying on the domain (see the README)."
  value       = var.domain != "" ? "https://${var.domain}" : "https://${openstack_networking_floatingip_v2.gateway.address}"
}
versions.tfHCL
terraform {
  required_version = ">= 1.6.0"

  required_providers {
    openstack = {
      source  = "terraform-provider-openstack/openstack"
      version = "~> 2.0"
    }
  }
}

provider "openstack" {}
terraform.tfvars.exampleHCL
# Required: SSH keypair must already exist in your project
key_name = "YOUR_KEY_NAME"

# Recommended: set a domain so the endpoint URL and TLS certificate use a stable
# name. Point its DNS A record at the floating IP and replace the first-boot
# self-signed certificate with a CA-issued one (see the README).
# domain = "gateway.example.com"

# litellm_version = "main-stable"
# flavor_name = "s1a.medium"
# image_name = "Ubuntu-24.04"
# app_name = "inference-gateway"
# volume_size = 20
# enable_db = true
# external_network = "PublicStatic"
# private_cidr = "10.60.0.0/24"

# Upstream provider API keys are NOT set here: they would land in Terraform
# state. After apply, SSH to the instance and populate /data/config/providers.env,
# then restart the stack. See the README, "Provider keys and the master key".
cloud-init/gateway.yaml.tftpl
#cloud-config
package_update: true
packages:
  - ca-certificates
  - curl
  - openssl

write_files:
  # First-boot bootstrap. Generates a self-signed TLS certificate and the
  # LiteLLM master key on the instance (and a Postgres password when the
  # database is enabled): none of them is shipped in this template. Writes a
  # starter LiteLLM config, an empty provider-keys file, and a docker compose
  # file, then brings the stack up. The master key is written to
  # /root/gateway-credentials (root-only); retrieve it over SSH.
  - path: /opt/gateway-bootstrap.sh
    permissions: "0755"
    content: |
      #!/usr/bin/env bash
      set -euo pipefail

      CRED_FILE=/root/gateway-credentials
      STACK_DIR=/opt/gateway
      CONFIG_DIR=/data/config
      CERT_DIR=/data/certs
      GW_HOST="${domain != "" ? domain : floating_ip}"

      # Idempotent across reboots: if already installed, just bring the stack up.
      if [ -f "$CRED_FILE" ]; then
        cd "$STACK_DIR"
        docker compose up -d
        exit 0
      fi

      mkdir -p "$CERT_DIR" "$CONFIG_DIR" "$STACK_DIR" /data/postgres

      # First-boot self-signed certificate. The proxy serves HTTPS only, so
      # clients need TLS from the first request. This cert lets them connect
      # after trusting it; replace it with a CA-issued certificate for anything
      # beyond a private network or a trial (see the README).
      openssl req -x509 -newkey rsa:4096 -nodes \
        -keyout "$CERT_DIR/gateway.key" \
        -out "$CERT_DIR/gateway.crt" \
        -days 825 \
        -subj "/CN=$GW_HOST" \
        -addext "subjectAltName=${domain != "" ? "DNS:${domain}" : "IP:${floating_ip}"}"
      chmod 600 "$CERT_DIR/gateway.key"

      # Generate the master key (clients authenticate to the gateway with it) and
      # the database password on the instance. Neither is shipped in the template.
      MASTER_KEY="sk-$(openssl rand -hex 24)"
      DB_PASSWORD="$(openssl rand -base64 24 | tr -dc 'A-Za-z0-9' | head -c 24)"

      # Starter routing config. model_name is what clients request; the
      # litellm_params.model is the provider/model it maps to. Provider keys come
      # from providers.env via os.environ, so no secret lives in this file.
      cat > "$CONFIG_DIR/litellm-config.yaml" <<'LITELLM_CONFIG'
      model_list:
        # Duplicate and edit one entry per upstream backend you route to (OpenAI,
        # Anthropic, a self-hosted vLLM/Ollama endpoint, or any OpenAI-compatible
        # GPU provider). The model inference runs on those backends; this VM only
        # routes requests to them.
        - model_name: gpt-4o-mini
          litellm_params:
            model: openai/gpt-4o-mini
            api_key: os.environ/OPENAI_API_KEY
        - model_name: my-openai-compatible-model
          litellm_params:
            model: openai/your-model-name
            api_base: os.environ/UPSTREAM_API_BASE
            api_key: os.environ/UPSTREAM_API_KEY

      litellm_settings:
        drop_params: true
        # Response caching cuts cost and latency on repeated prompts. The local
        # in-memory cache resets on restart; point cache_params at a Redis backend
        # for a cache that survives restarts and is shared across replicas.
        cache: true
        cache_params:
          type: local

      router_settings:
        # Fallback: when a model errors or rate-limits, route to a backup. Map a
        # primary model_name to an ordered list of backups, for example:
        # fallbacks: [{"gpt-4o-mini": ["my-openai-compatible-model"]}]
        routing_strategy: simple-shuffle

      general_settings:
        # Auth: every client request must carry this key as a bearer token. Mint
        # scoped virtual keys from the admin API when the database is enabled.
        master_key: os.environ/LITELLM_MASTER_KEY
%{ if enable_db }
        # Logging and spend tracking persist to this database.
        database_url: os.environ/DATABASE_URL
%{ endif }
      LITELLM_CONFIG

      # Provider keys live here, not in the template or in Terraform state. Fill
      # in the keys your config references, then: cd /opt/gateway && docker compose up -d
      cat > "$CONFIG_DIR/providers.env" <<'PROVIDERS_ENV'
      # Upstream provider API keys. Uncomment and set the ones your
      # litellm-config.yaml references via os.environ, then restart the stack.
      # OPENAI_API_KEY=
      # ANTHROPIC_API_KEY=
      # For a generic OpenAI-compatible backend (self-hosted vLLM, a GPU provider):
      # UPSTREAM_API_BASE=
      # UPSTREAM_API_KEY=
      PROVIDERS_ENV
      chmod 600 "$CONFIG_DIR/providers.env"

      # Docker compose for the gateway. The proxy terminates TLS itself and
      # publishes the one stable endpoint on 443; model inference happens on the
      # upstream backends, never on this VM.
      cat > "$STACK_DIR/docker-compose.yml" <<COMPOSE
      services:
        litellm:
          image: ghcr.io/berriai/litellm:${litellm_version}
          restart: unless-stopped
          ports:
            - "443:4000"
          env_file:
            - /data/config/providers.env
          environment:
            - LITELLM_MASTER_KEY=$MASTER_KEY
%{ if enable_db }
            - DATABASE_URL=postgresql://litellm:$DB_PASSWORD@db:5432/litellm
%{ endif }
          volumes:
            - /data/config/litellm-config.yaml:/app/config.yaml:ro
            - /data/certs:/data/certs:ro
          command: ["--config", "/app/config.yaml", "--port", "4000", "--ssl_certfile_path", "/data/certs/gateway.crt", "--ssl_keyfile_path", "/data/certs/gateway.key"]
%{ if enable_db }
          depends_on:
            - db
        db:
          image: postgres:16
          restart: unless-stopped
          environment:
            - POSTGRES_USER=litellm
            - POSTGRES_PASSWORD=$DB_PASSWORD
            - POSTGRES_DB=litellm
          volumes:
            - /data/postgres:/var/lib/postgresql/data
%{ endif }
      COMPOSE

      cd "$STACK_DIR"
      docker compose up -d

      umask 077
      cat > "$CRED_FILE" <<CRED
      Inference gateway master key (generated on first boot)
      endpoint:   ${domain != "" ? "https://${domain}" : "https://${floating_ip}"}
      master_key: $MASTER_KEY

      Clients send this as a bearer token. Mint scoped virtual keys from the admin
      API instead of sharing this one, then rotate it and delete this file.
      CRED
      chmod 600 "$CRED_FILE"

runcmd:
  - |
    set -e
    # The data volume attaches as /dev/sdb on this platform (not /dev/vdb).
    # Mount it at /data before the stack runs, so config, request logs, and the
    # database live on the resizable volume rather than the boot disk.
    DEV=/dev/sdb
    for i in $(seq 1 30); do [ -b "$DEV" ] && break; sleep 5; done
    if ! blkid "$DEV" >/dev/null 2>&1; then mkfs.ext4 -F -L gatewaydata "$DEV"; fi
    mkdir -p /data
    mount "$DEV" /data
    grep -q "$DEV" /etc/fstab || echo "$DEV /data ext4 defaults,nofail 0 2" >> /etc/fstab
    # Install Docker Engine (provides docker + the compose plugin).
    curl -fsSL https://get.docker.com | sh
    systemctl enable --now docker
    /opt/gateway-bootstrap.sh
README.mdMarkdown
# Inference gateway

Single compute instance running [LiteLLM](https://docs.litellm.ai), an open-source model router, on infrastructure you control. The gateway exposes one stable OpenAI-compatible endpoint over HTTPS and routes requests to any model backend you configure: hosted model APIs, a self-hosted vLLM or Ollama endpoint, or any OpenAI-compatible GPU provider. Clients point a single base URL at the gateway; auth, caching, fallback, and request logging live on the VM you own, and the heavy model inference runs on the upstream backends.


**Network class:** production — `external_network` defaults to `PublicStatic` for persisted floating IPs and multi-tier stacks; override with `PublicEphemeral` for ephemeral demos.

The instance provisions a private network, a floating IP, and a block volume mounted at `/data` so the config, request logs, and (by default) the key-and-spend database live on a resizable volume rather than the boot disk. cloud-init installs Docker, generates a self-signed TLS certificate and the master key, writes a starter routing config, and brings the stack up, all on first boot.

## Where this fits

This is the control-plane side of an AI workload: the always-on CPU layer that holds provider credentials, routing policy, and a single endpoint, while the model compute lives on backends that can be anywhere. It pairs with the data and runtime tracks in the IaC library: a [Redis cache](/resources/iac-templates/redis-cache) or [self-managed Postgres](/resources/iac-templates/self-managed-postgres) for app state, and a [containerized app](/resources/iac-templates/containerized-app) or [Next.js app](/resources/iac-templates/nextjs-app) for the agent or application that calls the gateway.

For a lighter, single-process router with no admin database, set `enable_db = false`; the gateway then authenticates with the master key only and keeps no request ledger.

## Prerequisites

- OpenTofu >= 1.6.0 or Terraform >= 1.6.0
- Quake AI account with OpenStack credentials
- An existing SSH keypair in your project (the value of `key_name` must match that keypair)
- API keys for the upstream model providers you intend to route to

## Resource baseline

The gateway is a lightweight router, not an inference host: LiteLLM, the TLS listener, and a small Postgres fit in 4 GiB. The default `s1a.medium` flavor (2 shared vCPU, 4 GiB RAM) handles steady proxy traffic; raise it if you front many backends or run high request concurrency. The boot disk is 30 GiB; config, logs, and the database live on the separate data volume (`volume_size`, default 20 GiB). No model weights are stored on this VM.

## Usage

1. Clone or copy this template directory
2. Copy `terraform.tfvars.example` to `terraform.tfvars` and set `key_name` (and `domain` if you have one)
3. Source your OpenStack credentials: `source openrc.sh`
4. Initialize: `tofu init`
5. Preview: `tofu plan`
6. Apply: `tofu apply`

cloud-init takes a few minutes on first boot to install Docker, pull the proxy image, and start the stack. The gateway then answers on `gateway_url` from the outputs, but it cannot reach any backend until you add provider keys (next section).

## Provider keys and the master key

No credential ships with this template, and provider keys never pass through Terraform state.

- **The master key** (how clients authenticate to the gateway) is generated on first boot and written to `/root/gateway-credentials` (mode 600). Retrieve it over SSH: `ssh user@<floating_ip> sudo cat /root/gateway-credentials`. Mint scoped virtual keys from the admin API instead of sharing the master key, then rotate it.
- **Upstream provider keys** are populated by you after apply, so they stay out of state and out of the repo:

```bash
ssh user@<floating_ip>
sudo nano /data/config/providers.env     # set OPENAI_API_KEY, UPSTREAM_API_KEY, etc.
cd /opt/gateway && sudo docker compose up -d   # restart to pick up the keys
```

The starter `/data/config/litellm-config.yaml` references those keys as `os.environ/OPENAI_API_KEY` and similar, so the secret stays in the env file, never in the config.

## Configuring routes, auth, caching, fallback, and logging

Edit `/data/config/litellm-config.yaml` on the instance and restart the stack to apply. The starter config documents each point:

- **Routes** (`model_list`): one entry per upstream model. `model_name` is what clients request; `litellm_params.model` is the provider/model it maps to. Add an `api_base` for self-hosted or third-party OpenAI-compatible endpoints.
- **Auth** (`general_settings.master_key`): every request carries the key as a bearer token. With the database enabled, mint per-team virtual keys with their own budgets from the admin API.
- **Caching** (`litellm_settings.cache`): a local in-memory cache is on by default; point `cache_params` at a Redis backend for a cache that survives restarts.
- **Fallback** (`router_settings`): map a primary `model_name` to an ordered list of backups so a model error or rate-limit routes to the next backend.
- **Logging**: the proxy logs to stdout (`docker compose logs litellm`); with the database enabled, per-request spend and usage persist to Postgres and surface in the admin UI.

## Ports and access

| Port | Purpose | Open to |
| --- | --- | --- |
| 22 | Host SSH for administration, setting provider keys, and retrieving the master key | `0.0.0.0/0` |
| 443 | The one stable gateway endpoint: the OpenAI-compatible API and admin UI over HTTPS | `0.0.0.0/0` |

The proxy serves HTTPS only; there is no plaintext port.

## Calling the gateway

The gateway speaks the OpenAI API. Point any OpenAI-compatible client at `gateway_url` with the master key (or a virtual key). Because the first-boot certificate is self-signed, a client either trusts the certificate or, for a quick test, skips verification:

```bash
curl https://HOST/v1/chat/completions \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "hello"}]}'
```

`HOST` is your domain or the floating IP; `gateway.crt` is the certificate copied from `/data/certs/gateway.crt`. The `model` value matches a `model_name` from your config.

## TLS for production

The first-boot certificate is self-signed, which is fine for a private network or a quick trial. For anything exposed to the internet:

- Set `domain`, point its DNS A record at `floating_ip`, and obtain a CA-issued certificate (for example with `certbot certonly`).
- Replace `/data/certs/gateway.crt` and `/data/certs/gateway.key` with the issued certificate and key, then `cd /opt/gateway && docker compose up -d` to apply.

With a CA-issued certificate, clients trust the gateway without the `--cacert` step above.

## Variables

| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `key_name` | string | yes | n/a | SSH keypair name (must already exist in your project) |
| `flavor_name` | string | no | `s1a.medium` | Instance size (4 GiB suits a router fronting several backends) |
| `image_name` | string | no | `Ubuntu-24.04` | Operating system image |
| `app_name` | string | no | `inference-gateway` | Display name prefix for resources |
| `litellm_version` | string | no | `main-stable` | LiteLLM proxy image tag; pin a versioned tag for reproducibility |
| `domain` | string | no | `""` | Public domain for the endpoint; empty uses the floating IP |
| `volume_size` | number | no | `20` | Block volume size in GiB, mounted at `/data` |
| `enable_db` | bool | no | `true` | Run Postgres for virtual keys, budgets, and spend tracking |
| `external_network` | string | no | `PublicStatic` | Persisted FIP / production default; override with `PublicEphemeral` for demos |
| `private_cidr` | string | no | `10.60.0.0/24` | CIDR for the private subnet |

## Outputs

| Name | Description |
| --- | --- |
| `floating_ip` | Public floating IP assigned to the instance |
| `private_ip` | Private IP address of the instance |
| `gateway_url` | The one stable gateway endpoint (domain when set, otherwise the floating IP) |
| `instance_id` | Compute instance ID |

## Scope

This is a single-VM gateway you operate, not a managed service. It is CPU-only and runs in one region; it routes to model and GPU backends that live elsewhere and makes no first-party GPU assumption. For high availability, run LiteLLM against external Postgres and Redis and place multiple instances behind a load balancer; this template provisions one node with its database and cache on the attached volume.

## Documentation

See also: [Redis cache template](/resources/iac-templates/redis-cache), [Containerized app template](/resources/iac-templates/containerized-app)
Resources, parameters, and variables
Provisions
Parameterized by
Variables
  • key_namerequired
  • flavor_name="s1a.medium"
  • image_name="Ubuntu-24.04"
  • app_name="inference-gateway"
  • litellm_version="main-stable"
  • domain=""
  • volume_size=20
  • enable_db=true
  • external_network="PublicStatic"
  • private_cidr="10.60.0.0/24"

Customize this pattern#

See also#

Usage Guidelines

The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.

For the full policy, see Usage Guidelines.

Last validated: 29.06.2026

Was this page helpful?