Model resource requirements
Reference for selecting a Quake AI instance flavor based on the language model you intend to run. All figures assume Ollama with 4-bit quantization (the default), CPU-only inference, and no other significant processes running on the instance.
How to read this table#
- RAM needed is the minimum to load and run the model. Add 2 GB headroom for the OS and Ollama overhead.
- Recommended flavor uses
{family}.{size}naming (for examplem2a.large,r2a.4xlarge) and targets comfortable headroom for inference at typical conversational throughput. See the flavor reference for vCPU and RAM per size. - Disk for weights is the storage required for the downloaded model weights. Ollama stores models in
/usr/share/ollama/.ollama/modelsby default. - Tokens/sec (est.) is a rough estimate on the recommended CPU flavor. GPU inference is 10–50× faster.
Small models (1–4 B parameters)#
Suitable for summarization, classification, simple Q&A, and code completion on constrained budgets.
| Model | Parameters | RAM needed | Recommended flavor | Disk for weights | Tokens/sec (est.) |
|---|---|---|---|---|---|
| Phi-4 Mini | 3.8 B | 4 GB | m2a.large | 2.5 GB | 20–35 |
| Phi-3 Mini | 3.8 B | 4 GB | m2a.large | 2.3 GB | 20–35 |
| Llama 3.2 3B | 3 B | 4 GB | m2a.large | 2.0 GB | 25–40 |
| Gemma 2 2B | 2 B | 3 GB | m2a.large | 1.6 GB | 30–50 |
| Qwen 2.5 1.5B | 1.5 B | 2 GB | s1a.medium | 1.0 GB | 40–60 |
Mid-size models (7–9 B parameters)#
The practical sweet spot for CPU inference: strong general capability with tolerable throughput.
| Model | Parameters | RAM needed | Recommended flavor | Disk for weights | Tokens/sec (est.) |
|---|---|---|---|---|---|
| Llama 3.1 8B | 8 B | 6 GB | m2a.xlarge | 4.7 GB | 10–20 |
| Mistral 7B | 7 B | 6 GB | m2a.xlarge | 4.1 GB | 12–22 |
| Qwen 2.5 7B | 7 B | 6 GB | m2a.xlarge | 4.4 GB | 12–20 |
| Gemma 2 9B | 9 B | 8 GB | m2a.xlarge | 5.5 GB | 8–15 |
| CodeLlama 7B | 7 B | 6 GB | m2a.xlarge | 3.8 GB | 12–22 |
Large models (13–14 B parameters)#
Better reasoning and instruction-following at the cost of throughput and RAM.
| Model | Parameters | RAM needed | Recommended flavor | Disk for weights | Tokens/sec (est.) |
|---|---|---|---|---|---|
| Llama 3.1 13B (Q4) | 13 B | 10 GB | m2a.2xlarge | 7.4 GB | 6–12 |
| CodeLlama 13B | 13 B | 10 GB | m2a.2xlarge | 7.4 GB | 6–12 |
| Phi-4 14B | 14 B | 10 GB | m2a.2xlarge | 9.1 GB | 5–10 |
Large models (30–70 B parameters)#
Frontier-class open-weight models. Require substantial RAM; practical throughput is low without GPU.
| Model | Parameters | RAM needed | Recommended flavor | Disk for weights | Tokens/sec (est.) |
|---|---|---|---|---|---|
| Llama 3.1 34B (Q4) | 34 B | 24 GB | r2a.4xlarge | 19 GB | 2–5 |
| Llama 3.1 70B (Q4) | 70 B | 48 GB | r2a.8xlarge | 40 GB | 1–3 |
| Qwen 2.5 72B (Q4) | 72 B | 48 GB | r2a.8xlarge | 44 GB | 1–3 |
Embedding models#
Embedding models are much lighter than generative models. They can run on the same Ollama instance as a generative model without significant performance impact.
| Model | Dimensions | RAM overhead | Disk | Notes |
|---|---|---|---|---|
nomic-embed-text | 768 | ~1 GB | 274 MB | Recommended default; strong retrieval quality |
mxbai-embed-large | 1024 | ~1 GB | 670 MB | Higher-dimensional vectors; better for large corpora |
all-minilm | 384 | < 1 GB | 46 MB | Smallest option; suitable for dev and prototyping |
bge-m3 | 1024 | ~1.5 GB | 1.2 GB | Multilingual; use when content is not English-only |
Quantization variants#
Ollama pulls 4-bit quantized models by default (Q4_K_M or similar). Higher quantization means better output quality at the cost of more RAM and slower inference.
| Quantization | Quality | RAM vs. Q4 | Pull tag example |
|---|---|---|---|
| Q2 | Noticeably degraded | ~50% | llama3.1:8b-instruct-q2_K |
| Q4 (default) | Good | 1× baseline | llama3.1:8b |
| Q5 | Better | ~1.25× | llama3.1:8b-instruct-q5_K_M |
| Q8 | Near-original | ~2× | llama3.1:8b-instruct-q8_0 |
Pull a specific quantization with ollama pull MODEL:TAG.
Storage configuration#
For large model collections, attach a block volume and set the Ollama model directory:
sudo mkdir -p /mnt/models
sudo chown ollama:ollama /mnt/models
sudo tee /etc/systemd/system/ollama.service.d/storage.conf <<EOF
[Service]
Environment="OLLAMA_MODELS=/mnt/models"
EOF
sudo systemctl daemon-reload
sudo systemctl restart ollamaSee How to attach a block volume for instructions on provisioning and mounting storage.
Inference hardware#
The Quake AI flavor catalog provides CPU-backed compute. Models in the 34–70 B parameter range are not practical for interactive inference on CPU; for those workloads, run inference on hardware outside Quake AI or against a managed inference provider.