Skip to content

Model Resource Requirements

Reference · Updated Sep 2026

Model resource requirements

Reference for selecting a Quake AI instance flavor based on the language model you intend to run. All figures assume Ollama with 4-bit quantization (the default), CPU-only inference, and no other significant processes running on the instance.

How to read this table#

  • RAM needed is the minimum to load and run the model. Add 2 GB headroom for the OS and Ollama overhead.
  • Recommended flavor uses {family}.{size} naming (for example m2a.large, r2a.4xlarge) and targets comfortable headroom for inference at typical conversational throughput. See the flavor reference for vCPU and RAM per size.
  • Disk for weights is the storage required for the downloaded model weights. Ollama stores models in /usr/share/ollama/.ollama/models by default.
  • Tokens/sec (est.) is a rough estimate on the recommended CPU flavor. GPU inference is 10–50× faster.

Small models (1–4 B parameters)#

Suitable for summarization, classification, simple Q&A, and code completion on constrained budgets.

ModelParametersRAM neededRecommended flavorDisk for weightsTokens/sec (est.)
Phi-4 Mini3.8 B4 GBm2a.large2.5 GB20–35
Phi-3 Mini3.8 B4 GBm2a.large2.3 GB20–35
Llama 3.2 3B3 B4 GBm2a.large2.0 GB25–40
Gemma 2 2B2 B3 GBm2a.large1.6 GB30–50
Qwen 2.5 1.5B1.5 B2 GBs1a.medium1.0 GB40–60

Mid-size models (7–9 B parameters)#

The practical sweet spot for CPU inference: strong general capability with tolerable throughput.

ModelParametersRAM neededRecommended flavorDisk for weightsTokens/sec (est.)
Llama 3.1 8B8 B6 GBm2a.xlarge4.7 GB10–20
Mistral 7B7 B6 GBm2a.xlarge4.1 GB12–22
Qwen 2.5 7B7 B6 GBm2a.xlarge4.4 GB12–20
Gemma 2 9B9 B8 GBm2a.xlarge5.5 GB8–15
CodeLlama 7B7 B6 GBm2a.xlarge3.8 GB12–22

Large models (13–14 B parameters)#

Better reasoning and instruction-following at the cost of throughput and RAM.

ModelParametersRAM neededRecommended flavorDisk for weightsTokens/sec (est.)
Llama 3.1 13B (Q4)13 B10 GBm2a.2xlarge7.4 GB6–12
CodeLlama 13B13 B10 GBm2a.2xlarge7.4 GB6–12
Phi-4 14B14 B10 GBm2a.2xlarge9.1 GB5–10

Large models (30–70 B parameters)#

Frontier-class open-weight models. Require substantial RAM; practical throughput is low without GPU.

ModelParametersRAM neededRecommended flavorDisk for weightsTokens/sec (est.)
Llama 3.1 34B (Q4)34 B24 GBr2a.4xlarge19 GB2–5
Llama 3.1 70B (Q4)70 B48 GBr2a.8xlarge40 GB1–3
Qwen 2.5 72B (Q4)72 B48 GBr2a.8xlarge44 GB1–3

Embedding models#

Embedding models are much lighter than generative models. They can run on the same Ollama instance as a generative model without significant performance impact.

ModelDimensionsRAM overheadDiskNotes
nomic-embed-text768~1 GB274 MBRecommended default; strong retrieval quality
mxbai-embed-large1024~1 GB670 MBHigher-dimensional vectors; better for large corpora
all-minilm384< 1 GB46 MBSmallest option; suitable for dev and prototyping
bge-m31024~1.5 GB1.2 GBMultilingual; use when content is not English-only

Quantization variants#

Ollama pulls 4-bit quantized models by default (Q4_K_M or similar). Higher quantization means better output quality at the cost of more RAM and slower inference.

QuantizationQualityRAM vs. Q4Pull tag example
Q2Noticeably degraded~50%llama3.1:8b-instruct-q2_K
Q4 (default)Good1× baselinellama3.1:8b
Q5Better~1.25×llama3.1:8b-instruct-q5_K_M
Q8Near-original~2×llama3.1:8b-instruct-q8_0

Pull a specific quantization with ollama pull MODEL:TAG.

Storage configuration#

For large model collections, attach a block volume and set the Ollama model directory:

bash
sudo mkdir -p /mnt/models
sudo chown ollama:ollama /mnt/models

sudo tee /etc/systemd/system/ollama.service.d/storage.conf <<EOF
[Service]
Environment="OLLAMA_MODELS=/mnt/models"
EOF

sudo systemctl daemon-reload
sudo systemctl restart ollama

See How to attach a block volume for instructions on provisioning and mounting storage.

Inference hardware#

The Quake AI flavor catalog provides CPU-backed compute. Models in the 34–70 B parameter range are not practical for interactive inference on CPU; for those workloads, run inference on hardware outside Quake AI or against a managed inference provider.

Was this page helpful?