Skip to content

Run a Local LLM with Ollama

Deployment · Updated May 2026

Coming from another cloud?

▸AWS·EC2 LLM

This Quake AI feature maps to AWS’s EC2 LLM.

▸DigitalOcean·Ollama Droplet

This Quake AI feature maps to DigitalOcean’s Ollama Droplet.

Run a local llm with ollama

Stand up Ollama on a Quake AI instance for local LLM inference. You run the model runtime yourself; this is a self-hosted inference endpoint you operate, not a managed model API.

Ollama packages models and their runtime into a single binary, handles model downloads, and exposes an OpenAI-compatible REST API on port 11434. CPU inference works for models up to about 8 billion parameters without a GPU.

Client appOllama API:11434Ollama processon Ubuntu instanceModel weightson local disk POST /api/generateload modeltokens
Click to zoom
Ollama on a Quake AI instance: client apps call the REST API; the process loads model weights from local disk

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on dedicated vCPU.

Starting template$127.00/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

m2a.xlarge

m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

m2a.xlarge

4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00

Compute + RAM rate basis

4 vCPU + 16 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$28.00/mo

Production on dedicated CPU

The headline estimate above; predictable steady-load performance.

$127.00/mo

Saves $99.00/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Prerequisites#

  • A Quake AI account with an active project
  • An SSH key pair added to your project
  • A security group or plan to configure one during instance creation

See How to create an instance if you need the console or CLI steps.

Choose an instance size#

Model performance scales directly with available RAM. Pick a flavor based on the model you intend to run.

Target modelParametersRAM neededRecommended flavor
Phi-3 Mini, Llama 3.2 3B1–3 B4 GBm2a.large (2 vCPU, 8 GiB)
Llama 3.1 8B, Mistral 7B7–8 B8 GBm2a.xlarge (4 vCPU, 16 GiB)
Llama 3.1 13B, CodeLlama 13B13 B16 GBm2a.2xlarge (8 vCPU, 32 GiB)
Llama 3.1 70B70 B48 GB+r2a.8xlarge (32 vCPU, 256 GiB) or larger

See model resource requirements for a full table including quantization variants.

This guide uses m2a.xlarge (4 vCPUs, 16 GiB RAM), which comfortably handles 7–8 B models and can run 3 B models with headroom.

Step 1: Create a VM#

Create a new instance:

  • Image: Ubuntu-24.04
  • Flavor: m2a.xlarge
  • Network: Attach to your project network
  • Security group: Allow inbound TCP on port 22 (SSH). You will open port 11434 in a later step if you need external API access
  • Key pair: Select your SSH key pair

Assign a floating IP after the instance launches.

If you need a walkthrough of instance creation, see How to create an instance.

Step 2: Install Ollama#

SSH into your VM:

bash
ssh ubuntu@YOUR_FLOATING_IP

Run the official installer:

bash
curl -fsSL https://ollama.ai/install.sh | sh

The installer places the ollama binary at /usr/local/bin/ollama, creates an ollama system user, and registers a systemd service that starts automatically. Verify it is running:

bash
systemctl status ollama

You should see active (running). The API is now listening on 127.0.0.1:11434.

Step 3: Pull and run a model#

Pull a model: llama3.2 is a good starting point on m2a.xlarge:

bash
ollama pull llama3.2

This downloads the model weights (approximately 2 GB). Progress is printed to the terminal.

Run an interactive session to verify inference works:

bash
ollama run llama3.2 "Explain Quake AI in one sentence."

You should receive a generated response within a few seconds. Type /bye to exit the interactive session.

Step 4: Verify the API#

Ollama exposes an OpenAI-compatible REST API locally:

bash
curl http://localhost:11434/api/generate \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "prompt": "What is retrieval-augmented generation?",
    "stream": false
  }'

The response includes a response field with the generated text. Set "stream": true for token-streaming responses.

List available models:

bash
curl http://localhost:11434/api/tags

Step 5: Allow external access (optional)#

By default, Ollama binds to 127.0.0.1 and is only reachable over SSH. If your application runs on a different VM, you have two options:

Option A: SSH tunnel (recommended for development):

From your local machine or application host, open a tunnel:

bash
ssh -L 11434:localhost:11434 ubuntu@YOUR_FLOATING_IP -N

Your application then connects to http://localhost:11434 through the tunnel. No firewall changes required.

Option B: Bind to all interfaces:

To expose the API on the VM's network interface, override the default bind address:

bash
sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf <<EOF
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
EOF
sudo systemctl daemon-reload
sudo systemctl restart ollama

The systemd override makes Ollama listen on 0.0.0.0:11434, but the security group from Step 1 only opens SSH. Add a security-group ingress rule for TCP 11434 from the source IPs that need to reach Ollama. Replace YOUR_CIDR with the CIDR range you want to allow (your office network, a peered VPC, or a single application VM's /32), and YOUR_SG_NAME with the security group attached to the VM:

bash
openstack security group rule create \
  --proto tcp --dst-port 11434 \
  --remote-ip YOUR_CIDR \
  YOUR_SG_NAME

Do not expose port 11434 to 0.0.0.0/0. Ollama has no built-in authentication.

Next steps#

Clean up#

Delete the instance from Compute > Instances when you no longer need the inference endpoint. Release the floating IP if you allocated one only for Ollama.

Troubleshooting#

ollama pull stalls or fails: Check your VM's outbound internet connectivity. The model registry is at registry.ollama.ai. If your network policy restricts outbound traffic, open port 443 to that host.

Inference is slow: Confirm the model fits within available RAM with free -h. If RAM is exhausted, the system will swap to disk, which is 10–50× slower. Switch to a smaller model or upgrade the instance flavor.

curl to port 11434 hangs from another host: The API binds to loopback by default. Follow Option A or B in Step 5 above.

Model outputs poor results: Try a larger quantization variant. Models like llama3.2:3b use 4-bit quantization by default. Pull llama3.2:3b-instruct-q8_0 for higher quality at the cost of more RAM.

Before this
Was this page helpful?