Add a Browser Interface to Ollama with Open WebUI
Add a browser interface to ollama with open webui
Stand up Open WebUI on the same Quake AI instance as Ollama. Open WebUI provides a browser chat UI with conversation history, model switching, and file uploads, backed by models on your instance.
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
m2a.large
m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
m2a.large
2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
2 vCPU + 8 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on dedicated CPU
The headline estimate above; predictable steady-load performance.
Saves $49.50/mo while you build on shared CPU.
Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Prerequisites#
- Ollama running with at least one model on your instance (Run a local LLM)
- Port 8080 open in your security group (restricted to your IP)
Step 1: Install Docker#
Install Docker if the Ollama host does not already have it:
curl -fsSL https://get.docker.com | sudo sh
sudo usermod -aG docker $USER
newgrp dockerVerify the install:
docker --versionStep 2: Run Open WebUI#
Deploy Open WebUI alongside Ollama using Docker:
docker run -d \
--name open-webui \
--restart unless-stopped \
--network host \
-v open-webui-data:/app/backend/data \
-e OLLAMA_BASE_URL=http://127.0.0.1:11434 \
ghcr.io/open-webui/open-webui:mainThe --network host flag lets the container reach Ollama directly on 127.0.0.1:11434 without extra networking configuration, and publishes the container's default port 8080 on the host. The security-group rule from the prerequisites covers external access without an additional port mapping. Data (conversations, settings, user accounts) persists in the open-webui-data Docker volume.
Step 3: create your admin account#
Open http://YOUR_FLOATING_IP:8080 in your browser.
On first launch, Open WebUI prompts you to create an admin account. This account controls access for all users. Choose a strong password. Subsequent users who sign up receive a "pending" role until an admin approves them.
Step 4: start a conversation#
Select a model from the dropdown at the top of the chat interface. All models you have pulled with ollama pull appear here automatically. Open WebUI polls Ollama's /api/tags endpoint on load, so new models appear without restarting anything.
Type a message and verify the response. If the interface shows a connection error:
docker logs open-webuiCommon cause: Ollama is not binding to 127.0.0.1:11434. Confirm with:
curl http://localhost:11434/api/tagsStep 5: enable file uploads for RAG (optional)#
Open WebUI includes a built-in RAG pipeline. To use it:
- Open Admin Panel → Documents
- Configure the embedding model: select an Ollama model (e.g.,
nomic-embed-text) - Upload files from the chat interface using the paperclip icon
Open WebUI chunks, embeds, and stores the documents in its local SQLite database. This is suitable for personal use and small teams. For larger-scale or shared RAG pipelines, use Qdrant and LangChain instead.
Managing users#
By default, Open WebUI requires all new users to be approved by an admin before they can chat. Navigate to Admin Panel → Users to manage roles and approve pending accounts.
For team deployments where you want all users to have immediate access, set the default role to user in Admin Panel → Settings → General.
Updating open WebUI#
Pull the latest image and recreate the container:
docker pull ghcr.io/open-webui/open-webui:main
docker stop open-webui
docker rm open-webui
docker run -d \
--name open-webui \
--restart unless-stopped \
--network host \
-v open-webui-data:/app/backend/data \
-e OLLAMA_BASE_URL=http://127.0.0.1:11434 \
ghcr.io/open-webui/open-webui:mainThe -v open-webui-data:/app/backend/data volume preserves all conversations, settings, and user accounts across updates.
Using Docker compose#
If you want to manage Ollama and Open WebUI together, a compose file keeps both services co-located:
services:
ollama:
image: ollama/ollama
restart: unless-stopped
network_mode: host
volumes:
- ollama-data:/root/.ollama
open-webui:
image: ghcr.io/open-webui/open-webui:main
restart: unless-stopped
network_mode: host
volumes:
- open-webui-data:/app/backend/data
environment:
- OLLAMA_BASE_URL=http://127.0.0.1:11434
depends_on:
- ollama
volumes:
ollama-data:
open-webui-data:Save as ~/ai-stack/docker-compose.yml and start with docker compose up -d. Pull models through the Open WebUI interface or from the terminal:
docker exec ollama ollama pull llama3.2Next steps#
- Run a local LLM with Ollama
- Build a RAG pipeline: production-grade retrieval with Qdrant
- Deploy Qdrant
Clean up#
Stop and remove the Open WebUI container when finished:
docker stop open-webui && docker rm open-webuiThe open-webui-data volume retains conversations until you remove it with docker volume rm open-webui-data.