This Quake AI feature maps to Google Cloud’s Vertex AI.
Deploy an inference gateway with OpenTofu
Stand up LiteLLM, an open-source model router, on one CPU VM using the validated OpenTofu templateinference-gateway. The gateway exposes a single OpenAI-compatible endpoint over HTTPS; model inference runs on the upstream backends you point at.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Gateway
s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps
$33.00/mo
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
$0.00
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
$0.00
Pricing data last validated: . For current rates, check quake.ai/pricing.
Click to zoom
Inference gateway topology: LiteLLM router with Postgres on a data volume, one floating IP on HTTPS, clients route to upstream backends through one endpoint
Copy terraform.tfvars.example to terraform.tfvars and set:
HCL
key_name = "YOUR_KEY_NAME"
Leave domain commented out for a self-signed certificate on the floating IP. Defaults for flavor, LiteLLM version, and volume size are documented on the Inference gateway reference page. Leave enable_db = true so the request-logging step has a database to read from.
From the template directory, run:
bash
tofu inittofu plantofu apply
Type yes when prompted. Provisioning takes a few minutes while cloud-init installs Docker, generates the certificate and master key, and starts the stack.