Skip to content
Solutions

AI inference and RAG pipelines

AI inference and RAG pipelines

Run the durable control plane for an inference or RAG workload on Quake AI: the serving API, the embedding and job workers, the data layer, the model router, and the credentials all live on Compute you operate, and the heavy model inference is a call out to a model backend you choose. Quake AI provides CPU-only Compute, Block Storage, Object Storage, and private networking. Quake AI flavors are AMD EPYC with no GPU option, which fits smaller models, CPU inference, and calling an external model API; see the compute FAQ for flavor details.

What this is for#

Teams building retrieval-augmented generation or self-hosted inference need a serving API, an embedding pipeline, a vector store, and the credentials that tie them to a model, all under their control. Quake AI provides CPU compute for serving and embedding, Block Storage for the database, and object storage for model artifacts and corpora; you operate the model runtime, the index, and the routing to whatever model backend serves the heavy inference. The outcome is a self-managed RAG and inference stack for smaller models, deployable from a validated OpenTofu template and its companion tutorial.

The control plane runs on Quake AI#

An inference or RAG workload is more than the model call. The parts that run continuously and hold state fit a flat-priced, always-on CPU tier, and they stay the same whichever model backend you point them at:

  • Serving and agent runtime. The long-running API or agent loop that accepts a request, retrieves context, and issues the model call.
  • Data gravity. Postgres with pgvector, Qdrant, Object Storage for model artifacts and corpora, and Redis for queues. The data a job reads and writes lives here.
  • Job dispatcher. A queue and workers that submit work to a model backend, poll, retry, and land results in Object Storage.
  • Model router. One stable endpoint that fans out to several model backends with auth, caching, fallback, and logging. The inference gateway template builds this tier.
  • Identity and secrets. Provider credentials and the spend policy stay on the instance you operate, inside your private network and security groups.
  • Observability. The monitoring stack watches latency, throughput, and cost from one place.

The heavy model inference is the one piece you place anywhere: serve a small model on the CPU instances above, or call an external model API (OpenAI, Anthropic, Google, or an open-weight host) from the same control plane. Re-pointing the model backend does not change the control plane around it.

OpenClaw, a one-click Console App, is this shape running today: a persistent AI agent on a Quake AI CPU VM that connects to any model provider, or to a local model through Ollama, with your API keys and conversation data staying on the VM you own.

Reference architecture#

API clientsQuake AIfull-stack-app templatevector variationInference API (CPU)Embedding workersPostgresObject Storage (models, corpora)pgvector index retrieve contextwrite embeddingsread corpora, modelssimilarity searchquery (HTTPS)
Click to zoom
RAG pipeline on Quake AI: the solid box is the full-stack-app base template (inference API, embedding workers, Postgres, and an Object Storage bucket for models and corpora); the dashed box is the vector variation that adds a pgvector index to the database. API clients query the inference API over HTTPS.

Download diagram: SVG, PNG, and PDF.

The pipeline is the full-stack app template plus a vector variation. Each tier maps to the diagram and to the template that builds it.

  1. API clients. Applications send queries to the inference API over HTTPS.

  2. Inference API. A serving process runs on Compute (CPU), accepts queries, retrieves context, and calls the model. The full-stack app template covers the API and database tier.

  3. Embedding workers. Compute workers turn source documents into embeddings, write them to the vector index, and read corpora and model artifacts from storage.

  4. Object Storage. Model weights, tokenizer files, and document corpora live in S3-compatible Object Storage, the bucket the template provisions.

  5. Vector variation. Embeddings are stored and searched in Postgres with pgvector on Block Storage. The vector variation adds the pgvector index to the database tier; the self-managed PostgreSQL template covers the same tier as a standalone stack.

Worked example: AI-assist for creator pipelines#

This worked example applies the same CPU-only stack to a creator media pipeline. It is a different shape of the same story, not a stronger CPU case: every component runs on the AMD EPYC compute described above, with no GPU. The components are whisper.cpp for transcription on a CPU worker, Piper or Coqui TTS for voice synthesis on a CPU worker, a small open-weight LLM served through Ollama (Llama 3.1 8B, Mistral 7B, or Qwen 2.5 7B) for caption and metadata generation, and Qdrant for the asset-search vector store. The self-hosted AI explainer documents all of these.

Three creator workflows run on this stack:

  1. Transcription. When an asset lands in Object Storage, a whisper.cpp worker transcribes a podcast or video back catalog into SRT or VTT captions.

  2. Caption and metadata generation. A small LLM generates a title, description, chapter markers, and social-post drafts from the transcript. The output writes back to the asset's metadata or to a webhook target.

  3. Asset search. Embedding workers index transcripts and visual descriptions in Qdrant. The creator's CMS queries Qdrant to find every clip that covers a given topic.

The CPU ceiling from the rest of this page applies without softening. This stack transcribes, generates captions, and indexes assets; it does not deliver GPU-quality voice cloning, image generation, or video generation. For the broader creator back end, see Creator businesses; for the audio pipeline this example often extends, see Podcast and audio publishing.

Services involved#

ServiceRole in this architectureDocs
ComputeInference API and embedding workers (CPU)Compute
Block StorageVector index and database volumesBlock Storage
Object StorageModel artifacts, tokenizers, and document corporaObject Storage
NetworkPrivate networks and security groups for the serving tierNetwork
Self-managed PostgresVector store and metadata you operateSelf-managed PostgreSQL template

Get started#

Estimate the cost#

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on a mix of shared and dedicated vCPU.

Starting template$119.30/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Web tier

s1a.small · 2 shared vCPU, 2 GiB RAM, 0.5 Gbps

$16.50/mo

App tier

s1a.medium · 4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00/mo

Database

m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

s1a.small

2 shared vCPU, 2 GiB RAM, 0.5 Gbps

$16.50

s1a.medium

4 shared vCPU, 4 GiB RAM, 0.5 Gbps

$33.00

m2a.large

2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00

Compute + RAM rate basis

8 vCPU + 14 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (110 GiB)

110 GiB at $0.08/GiB/mo

$8.80

Public IP (included)

1 included with the custom package

$0.00

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Configure your estimate

Check the add-ons you plan to deploy to build a monthly total. Nothing is selected to start, so the total below begins at the baseline.

Starting template

The required baseline, always included.

$119.30/mo
Your configured estimate$119.30/mo

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$69.80/mo

Production on the configured CPU

The headline estimate above; predictable steady-load performance.

$119.30/mo

Saves $49.50/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Migrating an existing AI inference / RAG service?#

Move an inference API or RAG stack that already runs on AWS EC2 and S3, Azure VMs and Blob, Google Compute Engine and Cloud Storage, or another hyperscaler. The outcome is the same serving and retrieval pipeline on Quake AI CPU compute with model artifacts, corpora, and vector indexes cut over in a controlled order.

Follow this cutover path. Each step links an existing migration page; this section composes those pages into a workload-shaped sequence rather than duplicating their steps.

  1. Map your source provider. Start with the concept-translation page for your current cloud: Coming from AWS, Coming from Azure, Coming from GCP, Coming from DigitalOcean, or Coming from Hetzner.

  2. Stand up the target shape on Quake AI. Pick the greenfield template that matches your serving layout: Full-stack app for an API plus database tier, or Self-managed PostgreSQL when vector-indexed Postgres is the primary data store.

  3. Move compute workloads. Rebuild or migrate inference and embedding workers with Migrate from EC2 (or the matching compute migration page for your source provider).

  4. Move object and artifact data. Sync model weights, document corpora, and embedding exports with Migrate from S3 (or the matching object migration page for your source provider).

  5. Cut over the inference endpoint. Stage the new API behind an inference gateway, API gateway, or private hostname. Run shadow traffic or canary requests against the Quake AI stack, then switch clients to the new base URL.

Workload-specific cutover callouts#

  • Model artifacts and corpora. Copy checkpoints, tokenizer files, and source documents before you switch production traffic. Validate checksums on large weight files; confirm the runtime loads the artifact from the Quake AI path you provisioned.
  • Vector store and index rebuild. Decide whether to transfer a vector-indexed Postgres or external index snapshot or rebuild embeddings from the migrated corpus. Rebuilds take longer but avoid version skew between embedding models and stored vectors.
  • CPU-only workloads. CPU inference fits smaller models and batch embedding jobs on Quake AI (compute FAQ). GPU-trained models that require GPU inference at production scale need accelerators outside this flavor catalog.
  • Endpoint cutover and rollback. Keep the source inference endpoint warm until latency and error rates stabilize on Quake AI. Roll back by switching clients back to the old base URL if quality checks or health probes fail.

Considerations and limits#

  • You operate serving and the index. Quake AI provides compute and storage; the model runtime, embedding pipeline, and vector index are yours under the shared responsibility model.
  • CPU-only compute. Compute is AMD EPYC with no GPU option (compute FAQ). CPU inference fits smaller models and batch embedding; GPU-required models at production scale need a different hosting path.
  • No managed database or vector service. Postgres with pgvector or another vector store runs on Compute and Block Storage you operate.
  • Flat egress. Quake AI applies a no-egress-fee policy for outbound transfer.
  • Three US regions. All current regions are in the United States.
  • Compliance posture. Quake AI holds SOC 2 Type I and Type II attestations and SOC 3. See Compliance and certifications for the platform scope.
Was this page helpful?