# Run a Local LLM with Ollama

Source: https://docs.quake.ai/resources/deployments/run-local-llm
Markdown: https://docs.quake.ai/resources/deployments/run-local-llm.md

---

# Run a local llm with ollama

Stand up [Ollama](https://ollama.ai) on a Quake AI instance for local LLM inference. You run the model runtime yourself; this is a self-hosted inference endpoint you operate, not a managed model API.

Ollama packages models and their runtime into a single binary, handles model downloads, and exposes an OpenAI-compatible REST API on port `11434`. CPU inference works for models up to about 8 billion parameters without a GPU.

<Figure size="md" caption="Ollama on a Quake AI instance: client apps call the REST API; the process loads model weights from local disk">

```d2
direction: right

client: Client app {shape: person}
api: Ollama API\n:11434
ollama: Ollama process\non Ubuntu instance
weights: Model weights\non local disk {shape: cylinder}

client -> api: POST /api/generate
api -> ollama
ollama -> weights: load model
ollama -> client: tokens
```

</Figure>

<PricingCompanion
  components={[
    { kind: "primitive", required: true, label: "Ollama host instance", vm: { flavor: "m2a.xlarge" } },
  ]}
/>

## Prerequisites

- A Quake AI account with an active project
- An SSH key pair added to your project
- A security group or plan to configure one during instance creation

See [How to create an instance](/docs/compute/how-to/create-instance) if you need the console or CLI steps.

## Choose an instance size

Model performance scales directly with available RAM. Pick a flavor based on the model you intend to run.

| Target model | Parameters | RAM needed | Recommended flavor |
|---|---|---|---|
| Phi-3 Mini, Llama 3.2 3B | 1–3 B | 4 GB | `m2a.large` (2 vCPU, 8 GiB) |
| Llama 3.1 8B, Mistral 7B | 7–8 B | 8 GB | `m2a.xlarge` (4 vCPU, 16 GiB) |
| Llama 3.1 13B, CodeLlama 13B | 13 B | 16 GB | `m2a.2xlarge` (8 vCPU, 32 GiB) |
| Llama 3.1 70B | 70 B | 48 GB+ | `r2a.8xlarge` (32 vCPU, 256 GiB) or larger |

See [model resource requirements](/reference/compute/model-resource-requirements) for a full table including quantization variants.

This guide uses `m2a.xlarge` (4 vCPUs, 16 GiB RAM), which comfortably handles 7–8 B models and can run 3 B models with headroom.

## Step 1: Create a VM

<CreateVmConsole flavor="m2a.xlarge" securityPorts="22 (SSH). You will open port 11434 in a later step if you need external API access" />

If you need a walkthrough of instance creation, see [How to create an instance](/docs/compute/how-to/create-instance).

## Step 2: Install Ollama

SSH into your VM:

```bash
ssh ubuntu@YOUR_FLOATING_IP
```

Run the official installer:

```bash
curl -fsSL https://ollama.ai/install.sh | sh
```

The installer places the `ollama` binary at `/usr/local/bin/ollama`, creates an `ollama` system user, and registers a systemd service that starts automatically. Verify it is running:

```bash
systemctl status ollama
```

You should see `active (running)`. The API is now listening on `127.0.0.1:11434`.

## Step 3: Pull and run a model

Pull a model: `llama3.2` is a good starting point on `m2a.xlarge`:

```bash
ollama pull llama3.2
```

This downloads the model weights (approximately 2 GB). Progress is printed to the terminal.

Run an interactive session to verify inference works:

```bash
ollama run llama3.2 "Explain Quake AI in one sentence."
```

You should receive a generated response within a few seconds. Type `/bye` to exit the interactive session.

## Step 4: Verify the API

Ollama exposes an OpenAI-compatible REST API locally:

```bash
curl http://localhost:11434/api/generate \
  -H "Content-Type: application/json" \
  -d '{
    "model": "llama3.2",
    "prompt": "What is retrieval-augmented generation?",
    "stream": false
  }'
```

The response includes a `response` field with the generated text. Set `"stream": true` for token-streaming responses.

List available models:

```bash
curl http://localhost:11434/api/tags
```

## Step 5: Allow external access (optional)

By default, Ollama binds to `127.0.0.1` and is only reachable over SSH. If your application runs on a different VM, you have two options:

**Option A: SSH tunnel (recommended for development):**

From your local machine or application host, open a tunnel:

```bash
ssh -L 11434:localhost:11434 ubuntu@YOUR_FLOATING_IP -N
```

Your application then connects to `http://localhost:11434` through the tunnel. No firewall changes required.

**Option B: Bind to all interfaces:**

To expose the API on the VM's network interface, override the default bind address:

```bash
sudo mkdir -p /etc/systemd/system/ollama.service.d
sudo tee /etc/systemd/system/ollama.service.d/override.conf <<EOF
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
EOF
sudo systemctl daemon-reload
sudo systemctl restart ollama
```

The systemd override makes Ollama listen on `0.0.0.0:11434`, but the security group from Step 1 only opens SSH. Add a security-group ingress rule for TCP 11434 from the source IPs that need to reach Ollama. Replace `YOUR_CIDR` with the CIDR range you want to allow (your office network, a peered VPC, or a single application VM's `/32`), and `YOUR_SG_NAME` with the security group attached to the VM:

```bash
openstack security group rule create \
  --proto tcp --dst-port 11434 \
  --remote-ip YOUR_CIDR \
  YOUR_SG_NAME
```

Do not expose port 11434 to `0.0.0.0/0`. Ollama has no built-in authentication.

## Next steps

- [Add Open WebUI](/resources/deployments/add-browser-ui-to-ollama): browser chat UI backed by your Ollama models
- [Deploy Qdrant](/resources/deployments/deploy-vector-database): vector storage for RAG pipelines
- [Build a RAG pipeline](/resources/deployments/deploy-rag-pipeline): combine Ollama inference with Qdrant retrieval

## Clean up

Delete the instance from **Compute** > **Instances** when you no longer need the inference endpoint. Release the floating IP if you allocated one only for Ollama.

## Troubleshooting

**`ollama pull` stalls or fails**: Check your VM's outbound internet connectivity. The model registry is at `registry.ollama.ai`. If your network policy restricts outbound traffic, open port 443 to that host.

**Inference is slow**: Confirm the model fits within available RAM with `free -h`. If RAM is exhausted, the system will swap to disk, which is 10–50× slower. Switch to a smaller model or upgrade the instance flavor.

**`curl` to port 11434 hangs from another host**: The API binds to loopback by default. Follow Option A or B in Step 5 above.

**Model outputs poor results**: Try a larger quantization variant. Models like `llama3.2:3b` use 4-bit quantization by default. Pull `llama3.2:3b-instruct-q8_0` for higher quality at the cost of more RAM.
