# Inference gateway

Source: https://docs.quake.ai/resources/iac-templates/inference-gateway
Markdown: https://docs.quake.ai/resources/iac-templates/inference-gateway.md

---

# Inference gateway

This pattern composes Compute, Network, and Block Storage into a self-hosted model router on infrastructure you control.

## What this template does

Provisions a single instance running [LiteLLM](https://docs.litellm.ai), an open-source model router, that exposes one stable OpenAI-compatible endpoint over HTTPS. The gateway routes requests to any model backend you configure: hosted model APIs, a self-hosted vLLM or Ollama endpoint, or any OpenAI-compatible GPU provider. Your clients point a single base URL at the gateway, and the auth, caching, fallback, and request logging run on a VM you own:

- Compute instance that runs the router in Docker, sized for proxy traffic rather than inference
- Private network, subnet, router, port, and security group; a floating IP for public access
- A block volume mounted at `/data`, so the routing config, request logs, and the key-and-spend database live on a volume you can grow rather than on the boot disk
- cloud-init installs Docker, generates a self-signed TLS certificate and the master key, writes a starter routing config, and brings the stack up on first boot

The model inference itself runs on the upstream backends, so this VM stays small. It is CPU-only, runs in one region, and routes to model and GPU backends that live elsewhere.

No credential ships with this template. The instance generates the master key on first boot and writes it to `/root/gateway-credentials` (readable only by root); retrieve it over SSH and rotate it after first use. Upstream provider keys are populated by you over SSH into `/data/config/providers.env`, so they stay out of Terraform state and out of the repo.

## With and without the database

By default the gateway runs an embedded Postgres for virtual API keys, per-key budgets, and spend tracking. Set `enable_db = false` for a single-process router that authenticates with the master key alone and keeps no request ledger. Use the database when you want to issue scoped keys to teams and track usage per key; the master-key-only mode suits a private gateway with one trusted caller.

## Parameters

| Parameter | Description | Default |
| --- | --- | --- |
| `key_name` | SSH keypair name (must already exist) | No default |
| `flavor_name` | Instance size (4 GiB suits a router fronting several backends) | `s1a.medium` |
| `image_name` | Operating system image | `Ubuntu-24.04` |
| `app_name` | Display name prefix for resources | `inference-gateway` |
| `litellm_version` | LiteLLM proxy image tag; pin a versioned tag for reproducibility | `main-stable` |
| `domain` | Public domain for the endpoint; empty uses the floating IP | `""` |
| `volume_size` | Block volume size in GiB, mounted at `/data` | `20` |
| `enable_db` | Run Postgres for virtual keys, budgets, and spend tracking | `true` |
| `external_network` | External network for floating IP allocation | `PublicStatic` |
| `private_cidr` | CIDR for the private subnet | `10.60.0.0/24` |

## Resource baseline

The gateway is a router, not an inference host: LiteLLM, the TLS listener, and a small Postgres fit in 4 GiB. The default `s1a.medium` flavor (2 shared vCPU, 4 GiB RAM) handles steady proxy traffic; raise it if you front many backends or run high request concurrency. The boot disk is 30 GiB; the config, logs, and database live on the separate data volume (`volume_size`, default 20 GiB). The VM stores no model weights.

## Ports and access

| Port | Purpose |
| --- | --- |
| 22 | Host SSH for administration, setting provider keys, and retrieving the master key |
| 443 | The one stable gateway endpoint: the OpenAI-compatible API and admin UI over HTTPS |

The proxy serves HTTPS only; there is no plaintext port. On first boot the certificate is self-signed, so a client trusts it by passing `/data/certs/gateway.crt` as the CA before the first request. For anything exposed to the internet, set `domain`, point its DNS A record at the floating IP, and replace the first-boot certificate with a CA-issued one. The README in the template directory covers both paths.

## How requests flow

The gateway speaks the OpenAI API. Point any OpenAI-compatible client at `gateway_url` with the master key (or a virtual key minted from the admin API). The `model` value in the request matches a `model_name` from your config, and the gateway forwards the call to the backend that `model_name` maps to:

```bash
curl https://HOST/v1/chat/completions \
  --cacert gateway.crt \
  -H "Authorization: Bearer $MASTER_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "gpt-4o-mini", "messages": [{"role": "user", "content": "hello"}]}'
```

You configure routes, auth, caching, fallback, and logging in `/data/config/litellm-config.yaml` on the instance. The starter config documents each point: `model_list` for the backends a request can map to, `general_settings.master_key` for auth, `litellm_settings.cache` for response caching, `router_settings` for model fallback, and the Postgres-backed spend log for per-request usage.

## When to use this pattern

Run your own model router on a VM you operate, so your apps and agents reach many model backends through one endpoint with one set of keys, instead of wiring each provider into each client. The control plane (the endpoint, the keys, the routing policy, the usage log) stays on infrastructure you own, while the model compute runs on whichever backends you point it at. It pairs with the runtime and data tracks in the IaC library: a [containerized app](/resources/iac-templates/containerized-app) or [Next.js app](/resources/iac-templates/nextjs-app) for the agent that calls the gateway, and a [Redis cache](/resources/iac-templates/redis-cache) or [self-managed Postgres](/resources/iac-templates/self-managed-postgres) for application state.

For a deeper treatment of running AI workloads on Quake AI, see [AI inference and RAG](/resources/solutions/ai-inference-rag).

## Estimated cost

<PricingCompanion
  components={[
    { kind: "template", slug: "inference-gateway", required: true },
  ]}
/>

## Template source

<TemplateSource slug="inference-gateway" />

<TemplateResourceMap template="inference-gateway" format="opentofu" />

## Customize this pattern

- [Customize a template's image and flavor](/docs/automation/how-to/customize-template-image-flavor)
- [Add a block volume to a template](/docs/automation/how-to/add-volume-to-template)
- [Parameterize a template with a tfvars file](/docs/automation/how-to/parameterize-template-tfvars)

## See also

- [Containerized app template](/resources/iac-templates/containerized-app)
- [Redis cache template](/resources/iac-templates/redis-cache)
