Run a Local LLM with Ollama
Coming from another cloud?
▸AWS·EC2 LLM
This Quake AI feature maps to AWS’s EC2 LLM.
▸DigitalOcean·Ollama Droplet
This Quake AI feature maps to DigitalOcean’s Ollama Droplet.
Run a local llm with ollama
Stand up Ollama on a Quake AI instance for local LLM inference. You run the model runtime yourself; this is a self-hosted inference endpoint you operate, not a managed model API.
Ollama packages models and their runtime into a single binary, handles model downloads, and exposes an OpenAI-compatible REST API on port 11434. CPU inference works for models up to about 8 billion parameters without a GPU.
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
m2a.xlarge
m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
m2a.xlarge
4 dedicated vCPU, 16 GiB RAM, 1 Gbps
Compute + RAM rate basis
4 vCPU + 16 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on dedicated CPU
The headline estimate above; predictable steady-load performance.
Saves $99.00/mo while you build on shared CPU.
Shared flavors carry less RAM (m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Prerequisites#
- A Quake AI account with an active project
- An SSH key pair added to your project
- A security group or plan to configure one during instance creation
See How to create an instance if you need the console or CLI steps.
Choose an instance size#
Model performance scales directly with available RAM. Pick a flavor based on the model you intend to run.
| Target model | Parameters | RAM needed | Recommended flavor |
|---|---|---|---|
| Phi-3 Mini, Llama 3.2 3B | 1–3 B | 4 GB | m2a.large (2 vCPU, 8 GiB) |
| Llama 3.1 8B, Mistral 7B | 7–8 B | 8 GB | m2a.xlarge (4 vCPU, 16 GiB) |
| Llama 3.1 13B, CodeLlama 13B | 13 B | 16 GB | m2a.2xlarge (8 vCPU, 32 GiB) |
| Llama 3.1 70B | 70 B | 48 GB+ | r2a.8xlarge (32 vCPU, 256 GiB) or larger |
See model resource requirements for a full table including quantization variants.
This guide uses m2a.xlarge (4 vCPUs, 16 GiB RAM), which comfortably handles 7–8 B models and can run 3 B models with headroom.
Step 1: Create a VM#
Create a new instance:
- Image: Ubuntu-24.04
- Flavor:
m2a.xlarge - Network: Attach to your project network
- Security group: Allow inbound TCP on port 22 (SSH). You will open port 11434 in a later step if you need external API access
- Key pair: Select your SSH key pair
Assign a floating IP after the instance launches.
If you need a walkthrough of instance creation, see How to create an instance.
Step 2: Install Ollama#
SSH into your VM:
ssh ubuntu@YOUR_FLOATING_IPRun the official installer:
curl -fsSL https://ollama.ai/install.sh | shThe installer places the ollama binary at /usr/local/bin/ollama, creates an ollama system user, and registers a systemd service that starts automatically. Verify it is running:
systemctl status ollamaYou should see active (running). The API is now listening on 127.0.0.1:11434.
Step 3: Pull and run a model#
Pull a model: llama3.2 is a good starting point on m2a.xlarge:
ollama pull llama3.2This downloads the model weights (approximately 2 GB). Progress is printed to the terminal.
Run an interactive session to verify inference works:
ollama run llama3.2 "Explain Quake AI in one sentence."You should receive a generated response within a few seconds. Type /bye to exit the interactive session.
Step 4: Verify the API#
Ollama exposes an OpenAI-compatible REST API locally:
curl http://localhost:11434/api/generate \
-H "Content-Type: application/json" \
-d '{
"model": "llama3.2",
"prompt": "What is retrieval-augmented generation?",
"stream": false
}'The response includes a response field with the generated text. Set "stream": true for token-streaming responses.
List available models:
curl http://localhost:11434/api/tagsStep 5: Allow external access (optional)#
By default, Ollama binds to 127.0.0.1 and is only reachable over SSH. If your application runs on a different VM, you have two options:
Option A: SSH tunnel (recommended for development):
From your local machine or application host, open a tunnel:
ssh -L 11434:localhost:11434 ubuntu@YOUR_FLOATING_IP -NYour application then connects to http://localhost:11434 through the tunnel. No firewall changes required.
Option B: Bind to all interfaces:
To expose the API on the VM's network interface, override the default bind address:
sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf <<EOF
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
EOF
sudo systemctl daemon-reload
sudo systemctl restart ollamaThe systemd override makes Ollama listen on 0.0.0.0:11434, but the security group from Step 1 only opens SSH. Add a security-group ingress rule for TCP 11434 from the source IPs that need to reach Ollama. Replace YOUR_CIDR with the CIDR range you want to allow (your office network, a peered VPC, or a single application VM's /32), and YOUR_SG_NAME with the security group attached to the VM:
openstack security group rule create \
--proto tcp --dst-port 11434 \
--remote-ip YOUR_CIDR \
YOUR_SG_NAMEDo not expose port 11434 to 0.0.0.0/0. Ollama has no built-in authentication.
Next steps#
- Add Open WebUI: browser chat UI backed by your Ollama models
- Deploy Qdrant: vector storage for RAG pipelines
- Build a RAG pipeline: combine Ollama inference with Qdrant retrieval
Clean up#
Delete the instance from Compute > Instances when you no longer need the inference endpoint. Release the floating IP if you allocated one only for Ollama.
Troubleshooting#
ollama pull stalls or fails: Check your VM's outbound internet connectivity. The model registry is at registry.ollama.ai. If your network policy restricts outbound traffic, open port 443 to that host.
Inference is slow: Confirm the model fits within available RAM with free -h. If RAM is exhausted, the system will swap to disk, which is 10–50× slower. Switch to a smaller model or upgrade the instance flavor.
curl to port 11434 hangs from another host: The API binds to loopback by default. Follow Option A or B in Step 5 above.
Model outputs poor results: Try a larger quantization variant. Models like llama3.2:3b use 4-bit quantization by default. Pull llama3.2:3b-instruct-q8_0 for higher quality at the cost of more RAM.