Skip to content

Deploy a Vector Database with Qdrant

Deployment · Updated May 2026

Coming from another cloud?

▸AWS·Opensearch Vector

This Quake AI feature maps to AWS’s Opensearch Vector.

▸DigitalOcean·Managed Vector DB

This Quake AI feature maps to DigitalOcean’s Managed Vector DB.

Deploy a vector database with qdrant

Stand up Qdrant on a Quake AI instance for vector search, embedding storage, and retrieval-augmented generation. You run the database yourself; this is a self-hosted vector store you operate, not a managed vector service.

Pair the instance with Ollama on the same host or a peer VM for a self-hosted RAG stack.

Your appQdrant API:6333Qdrant containeron Ubuntu instanceVector storage(volume) upsert / searchpersistmatched vectors
Click to zoom
Qdrant on a Quake AI instance: applications upsert and search vectors; data persists on local or block storage

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on dedicated vCPU.

Starting template$63.40/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Qdrant

m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

m2a.large

2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00

Compute + RAM rate basis

2 vCPU + 8 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (30 GiB)

30 GiB at $0.08/GiB/mo

$2.40

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$13.90/mo

Production on dedicated CPU

The headline estimate above; predictable steady-load performance.

$63.40/mo

Saves $49.50/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Prerequisites#

  • A Quake AI account with an active project
  • An SSH key pair added to your project

Choose an instance size#

Qdrant is memory-efficient. For most development and small-to-medium production workloads, m2a.large (2 vCPUs, 8 GiB RAM) is sufficient. Size up based on the number of vectors you plan to store:

Approximate vector countRecommended flavorNotes
Up to 1 millionm2a.large (2 vCPU, 8 GiB)Comfortable for dev and small prod
1–10 millionm2a.xlarge (4 vCPU, 16 GiB)Add a block volume for storage
10 million+m2a.2xlarge (8 vCPU, 32 GiB) or largerConsider dedicated storage volume

Vector dimensions affect memory: 1 million vectors at 768 dimensions uses about 3 GB of RAM. Use the Qdrant capacity calculator for precise estimates.

Step 1: Create a VM#

Create a new instance:

  • Image: Ubuntu-24.04
  • Flavor: m2a.large
  • Network: Attach to your project network
  • Security group: Allow inbound TCP on port 22 (SSH). Port 6333 should stay restricted; see Step 4
  • Key pair: Select your SSH key pair

Assign a floating IP after the instance launches.

Step 2: Install Docker#

SSH into your VM:

ssh ubuntu@YOUR_FLOATING_IP

Install Docker and Docker Compose:

curl -fsSL https://get.docker.com | sh

Add your user to the Docker group so you can run commands without sudo:

sudo usermod -aG docker ubuntu
newgrp docker

Group membership changes take effect on your next login. The newgrp docker line above activates the group in your current shell. If you skip it or open a new shell, log out and SSH back in before you run docker commands.

Confirm the install with a command that needs the daemon socket:

docker run --rm hello-world

A daemon-less check such as docker --version passes even when the group is not active yet, so it hides this gotcha.

Verify the installation:

bash
docker --version

Step 3: Run Qdrant#

Start Qdrant with persistent storage:

bash
mkdir -p ~/qdrant-data

docker run -d \
  --name qdrant \
  --restart unless-stopped \
  -p 127.0.0.1:6333:6333 \
  -p 127.0.0.1:6334:6334 \
  -v ~/qdrant-data:/qdrant/storage:z \
  qdrant/qdrant

The -p 127.0.0.1:6333:6333 binding keeps the API on loopback. Data in ~/qdrant-data persists across container restarts.

Verify Qdrant is running:

bash
curl http://localhost:6333/

Expected response:

JSON
{"title":"qdrant - vector search engine","version":"..."}

View the built-in dashboard (from your local machine via SSH tunnel; see Step 4):

http://localhost:6333/dashboard

Step 4: Configure access#

For applications on the same VM, use http://localhost:6333 directly; no configuration needed.

For applications on other VMs in the same project network, connect using the VM's private IP. Update the Docker binding to listen on the private interface:

bash
docker stop qdrant && docker rm qdrant

docker run -d \
  --name qdrant \
  --restart unless-stopped \
  -p PRIVATE_IP:6333:6333 \
  -p PRIVATE_IP:6334:6334 \
  -v ~/qdrant-data:/qdrant/storage:z \
  qdrant/qdrant

Replace PRIVATE_IP with the VM's private network address (visible in the console or via ip a). Update your project's security group to allow TCP on port 6333 from your application VMs' private CIDRs only.

For access from your local machine, use an SSH tunnel:

bash
ssh -L 6333:localhost:6333 ubuntu@YOUR_FLOATING_IP -N

Then open http://localhost:6333/dashboard in your browser.

Step 5: Create a collection and insert vectors#

Test the API with a small example using Python.

Ubuntu 24.04 enforces PEP 668, which blocks pip install against the system Python. Create a virtual environment first:

bash
sudo apt-get update
sudo apt-get install -y python3-venv

python3 -m venv ~/qdrant-venv
source ~/qdrant-venv/bin/activate

Install the Qdrant client inside the venv:

bash
pip install qdrant-client

Create a collection, insert sample vectors, and run a query. The query call returns a QueryResponse whose .points attribute is the list of matched points.

Python
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct

client = QdrantClient(url="http://localhost:6333")

# Create a collection with 384-dimensional vectors (all-MiniLM-L6-v2 output size)
client.create_collection(
    collection_name="documents",
    vectors_config=VectorParams(size=384, distance=Distance.COSINE),
)

# Insert sample vectors
client.upsert(
    collection_name="documents",
    points=[
        PointStruct(id=1, vector=[0.1] * 384, payload={"text": "Quake AI compute documentation"}),
        PointStruct(id=2, vector=[0.2] * 384, payload={"text": "OpenStack networking guide"}),
    ],
)

# Search
results = client.query_points(
    collection_name="documents",
    query=[0.15] * 384,
    limit=2,
).points
for r in results:
    print(r.id, r.score, r.payload)

For a complete RAG implementation using Ollama for embeddings, see Build a RAG pipeline.

Next steps#

  • Run a local LLM: add Ollama on the same VM or a separate instance for retrieval-augmented generation
  • Build a RAG pipeline: ingest documents, embed with Ollama, store in Qdrant, and query the result

Clean up#

Stop and remove the container, then delete the instance when finished:

bash
docker stop qdrant && docker rm qdrant

Troubleshooting#

Container exits immediately: Run docker logs qdrant to see the error. Storage permission issues (chmod 777 ~/qdrant-data or a :z label on the volume mount) are the most common cause on SELinux-enabled systems.

Out of memory on insert: Qdrant loads vectors into RAM for indexing. If you're inserting large batches, increase the instance flavor or insert in smaller chunks.

Dashboard is blank: The dashboard is a single-page app served from /dashboard. If the page loads but shows no data, check that collections exist with curl http://localhost:6333/collections.

Before this
Was this page helpful?