Self-Hosted AI on Cloud Infrastructure
Self-hosted AI on cloud infrastructure
Self-hosted AI means running language models on infrastructure you control. Your prompts and completions stay on your network. You pay a fixed compute cost (not per token), and you can run open-weight models. This page covers the components required, the trade-offs versus managed APIs, and common deployment patterns on Quake AI.
What "self-hosted AI" means in practice#
Running a language model on a VM requires more than downloading a model file and executing it. A functional self-hosted AI stack has these layers:
Inference runtime: Software that loads a model and serves it over an API. Ollama is the most accessible option for CPU inference; it handles model downloads, quantization, and exposes an OpenAI-compatible REST endpoint. Other options include llama.cpp (lower-level), LM Studio (desktop), and vLLM (GPU-optimized, higher throughput).
Model weights: The actual parameters of the model, typically distributed in GGUF format for CPU inference or safetensors for GPU. Common open-weight models: Llama 3.1/3.2 (Meta), Mistral 7B (Mistral AI), Phi-3/Phi-4 (Microsoft), Qwen 2.5 (Alibaba), Gemma 2 (Google). Most are available through Ollama's model registry.
Embedding model: A separate, smaller model that converts text into dense vectors for semantic search and retrieval. Common options: nomic-embed-text (via Ollama), all-MiniLM-L6-v2, bge-small-en. Embeddings are cheaper to run than generative models and can share the same Ollama instance.
Vector database: Stores the embeddings and enables similarity search. Qdrant is a common choice: open-source and simple to run in Docker. Alternatives include Chroma (lighter, embedded mode) and Weaviate.
Application layer: Your code that ties inference and retrieval together. LangChain and LlamaIndex are widely used Python orchestration libraries. For no-code workflows, Flowise and n8n provide visual builders.
The case for self-hosted AI on IaaS#
Data sovereignty. Prompts, completions, and documents stay on your infrastructure. For regulated industries, confidential research, or internal tooling with sensitive data, that constraint is often non-negotiable.
Cost predictability. Quake AI bills for CPU and RAM by the hour, regardless of inference volume. At high request volumes, thousands of daily calls, the fixed VM cost can be lower than per-token pricing from a managed API.
Model choice. Managed APIs expose a curated set of models. Self-hosting gives access to open-weight models, including domain-specific fine-tunes and older versions.
Latency control. A VM in a region close to your users delivers lower first-token latency than a shared API endpoint with variable queue depth.
No rate limits. Self-hosted inference is constrained only by the VM's compute capacity, which you can increase by resizing or adding instances.
The case against#
Operational overhead. You own the runtime: model updates, security patches, and uptime monitoring. Managed APIs handle this automatically.
CPU inference is slower than GPU. Running a 7 B parameter model on a CPU produces 10–30 tokens per second. GPT-4-class models at managed providers return hundreds of tokens per second from GPU clusters. Latency-sensitive, high-throughput applications typically need GPU inference or a managed API.
Model quality ceiling. The best open-weight models (Llama 3.1 70B, Qwen 2.5 72B) are competitive with GPT-4-class models for many tasks, but frontier models like Claude 3.5 Sonnet and GPT-4o are not open-weight and cannot be self-hosted.
Cold engineering cost. Standing up an inference runtime, embedding model, and vector database requires more initial work than calling a managed API with an HTTP client.
Common architectures on Quake AI#
Single-VM stack. Ollama and Qdrant on the same instance. Suitable for development, small teams, and low-traffic applications. Use m2a.xlarge for models up to 8 B parameters; m2a.2xlarge for 13 B models.
Split inference and storage. Ollama on one VM, Qdrant on a second VM with an attached block volume. Lets inference and retrieval scale independently.
Agent host + inference backend. An agent framework (Flowise, n8n, OpenClaw) on one VM, Ollama on another. The agent VM can be small (m2a.large); the inference VM sizes to the model. See Run a local LLM with Ollama for configuration details.
Kubernetes. For multi-tenant or high-availability deployments, run Ollama and Qdrant as Kubernetes Deployments on a Quake AI Kubernetes cluster. Horizontal Pod Autoscaler can scale inference replicas by CPU pressure.
Inference performance ceiling#
The Quake AI flavor catalog provides CPU-backed compute. The practical limit for responsive CPU inference of language models is approximately 7–8 billion parameters on m2a.xlarge or similar general-purpose flavors. Workloads that need GPU-accelerated inference (larger models, sub-second latency, high throughput) run on hardware outside Quake AI or against a managed inference provider.
Storage considerations#
Model weights are large. Common sizes:
- 1–3 B parameter models: 1–2 GB
- 7–8 B parameter models: 4–5 GB
- 13 B parameter models: 8–10 GB
- 70 B parameter models: 40–45 GB
Ollama stores downloaded models in /usr/share/ollama/.ollama/models by default. For large model collections or shared model caches across VMs, attach a block volume and configure Ollama's model directory with OLLAMA_MODELS=/your/volume/path.
Object storage (S3-compatible) is suitable for cold model archives and dataset storage, but Ollama reads from local disk. Copy weights to block storage before loading.