Deploy a Vector Database with Qdrant
Coming from another cloud?
▸AWS·Opensearch Vector
This Quake AI feature maps to AWS’s Opensearch Vector.
▸DigitalOcean·Managed Vector DB
This Quake AI feature maps to DigitalOcean’s Managed Vector DB.
Deploy a vector database with qdrant
Stand up Qdrant on a Quake AI instance for vector search, embedding storage, and retrieval-augmented generation. You run the database yourself; this is a self-hosted vector store you operate, not a managed vector service.
Pair the instance with Ollama on the same host or a peer VM for a self-hosted RAG stack.
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Qdrant
m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
m2a.large
2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
2 vCPU + 8 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (30 GiB)
30 GiB at $0.08/GiB/mo
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on dedicated CPU
The headline estimate above; predictable steady-load performance.
Saves $49.50/mo while you build on shared CPU.
Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Prerequisites#
- A Quake AI account with an active project
- An SSH key pair added to your project
Choose an instance size#
Qdrant is memory-efficient. For most development and small-to-medium production workloads, m2a.large (2 vCPUs, 8 GiB RAM) is sufficient. Size up based on the number of vectors you plan to store:
| Approximate vector count | Recommended flavor | Notes |
|---|---|---|
| Up to 1 million | m2a.large (2 vCPU, 8 GiB) | Comfortable for dev and small prod |
| 1–10 million | m2a.xlarge (4 vCPU, 16 GiB) | Add a block volume for storage |
| 10 million+ | m2a.2xlarge (8 vCPU, 32 GiB) or larger | Consider dedicated storage volume |
Vector dimensions affect memory: 1 million vectors at 768 dimensions uses about 3 GB of RAM. Use the Qdrant capacity calculator
Step 1: Create a VM#
Create a new instance:
- Image: Ubuntu-24.04
- Flavor:
m2a.large - Network: Attach to your project network
- Security group: Allow inbound TCP on port 22 (SSH). Port 6333 should stay restricted; see Step 4
- Key pair: Select your SSH key pair
Assign a floating IP after the instance launches.
Step 2: Install Docker#
SSH into your VM:
ssh ubuntu@YOUR_FLOATING_IPInstall Docker and Docker Compose:
curl -fsSL https://get.docker.com | shAdd your user to the Docker group so you can run commands without sudo:
sudo usermod -aG docker ubuntu
newgrp dockerGroup membership changes take effect on your next login. The newgrp docker line above activates the group in your current shell. If you skip it or open a new shell, log out and SSH back in before you run docker commands.
Confirm the install with a command that needs the daemon socket:
docker run --rm hello-worldA daemon-less check such as docker --version passes even when the group is not active yet, so it hides this gotcha.
Verify the installation:
docker --versionStep 3: Run Qdrant#
Start Qdrant with persistent storage:
mkdir -p ~/qdrant-data
docker run -d \
--name qdrant \
--restart unless-stopped \
-p 127.0.0.1:6333:6333 \
-p 127.0.0.1:6334:6334 \
-v ~/qdrant-data:/qdrant/storage:z \
qdrant/qdrantThe -p 127.0.0.1:6333:6333 binding keeps the API on loopback. Data in ~/qdrant-data persists across container restarts.
Verify Qdrant is running:
curl http://localhost:6333/Expected response:
{"title":"qdrant - vector search engine","version":"..."}View the built-in dashboard (from your local machine via SSH tunnel; see Step 4):
http://localhost:6333/dashboardStep 4: Configure access#
For applications on the same VM, use http://localhost:6333 directly; no configuration needed.
For applications on other VMs in the same project network, connect using the VM's private IP. Update the Docker binding to listen on the private interface:
docker stop qdrant && docker rm qdrant
docker run -d \
--name qdrant \
--restart unless-stopped \
-p PRIVATE_IP:6333:6333 \
-p PRIVATE_IP:6334:6334 \
-v ~/qdrant-data:/qdrant/storage:z \
qdrant/qdrantReplace PRIVATE_IP with the VM's private network address (visible in the console or via ip a). Update your project's security group to allow TCP on port 6333 from your application VMs' private CIDRs only.
For access from your local machine, use an SSH tunnel:
ssh -L 6333:localhost:6333 ubuntu@YOUR_FLOATING_IP -NThen open http://localhost:6333/dashboard in your browser.
Step 5: Create a collection and insert vectors#
Test the API with a small example using Python.
Ubuntu 24.04 enforces PEP 668pip install against the system Python. Create a virtual environment first:
sudo apt-get update
sudo apt-get install -y python3-venv
python3 -m venv ~/qdrant-venv
source ~/qdrant-venv/bin/activateInstall the Qdrant client inside the venv:
pip install qdrant-clientCreate a collection, insert sample vectors, and run a query. The query call returns a QueryResponse whose .points attribute is the list of matched points.
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct
client = QdrantClient(url="http://localhost:6333")
# Create a collection with 384-dimensional vectors (all-MiniLM-L6-v2 output size)
client.create_collection(
collection_name="documents",
vectors_config=VectorParams(size=384, distance=Distance.COSINE),
)
# Insert sample vectors
client.upsert(
collection_name="documents",
points=[
PointStruct(id=1, vector=[0.1] * 384, payload={"text": "Quake AI compute documentation"}),
PointStruct(id=2, vector=[0.2] * 384, payload={"text": "OpenStack networking guide"}),
],
)
# Search
results = client.query_points(
collection_name="documents",
query=[0.15] * 384,
limit=2,
).points
for r in results:
print(r.id, r.score, r.payload)For a complete RAG implementation using Ollama for embeddings, see Build a RAG pipeline.
Next steps#
- Run a local LLM: add Ollama on the same VM or a separate instance for retrieval-augmented generation
- Build a RAG pipeline: ingest documents, embed with Ollama, store in Qdrant, and query the result
Clean up#
Stop and remove the container, then delete the instance when finished:
docker stop qdrant && docker rm qdrantTroubleshooting#
Container exits immediately: Run docker logs qdrant to see the error. Storage permission issues (chmod 777 ~/qdrant-data or a :z label on the volume mount) are the most common cause on SELinux-enabled systems.
Out of memory on insert: Qdrant loads vectors into RAM for indexing. If you're inserting large batches, increase the instance flavor or insert in smaller chunks.
Dashboard is blank: The dashboard is a single-page app served from /dashboard. If the page loads but shows no data, check that collections exist with curl http://localhost:6333/collections.