# Deploy a Vector Database with Qdrant

Source: https://docs.quake.ai/resources/deployments/deploy-vector-database
Markdown: https://docs.quake.ai/resources/deployments/deploy-vector-database.md

---

# Deploy a vector database with qdrant

Stand up [Qdrant](https://qdrant.tech) on a Quake AI instance for vector search, embedding storage, and retrieval-augmented generation. You run the database yourself; this is a self-hosted vector store you operate, not a managed vector service.

Pair the instance with [Ollama](/resources/deployments/run-local-llm) on the same host or a peer VM for a self-hosted RAG stack.

<Figure size="md" caption="Qdrant on a Quake AI instance: applications upsert and search vectors; data persists on local or block storage">

```d2
direction: right

app: Your app {shape: person}
api: Qdrant API\n:6333
qdrant: Qdrant container\non Ubuntu instance
storage: Vector storage\n(volume) {shape: cylinder}

app -> api: upsert / search
api -> qdrant
qdrant -> storage: persist
qdrant -> app: matched vectors
```

</Figure>

<PricingCompanion components={[{ kind: "template", slug: "qdrant", required: true }]} />

## Prerequisites

- A Quake AI account with an active project
- An SSH key pair added to your project

## Choose an instance size

Qdrant is memory-efficient. For most development and small-to-medium production workloads, `m2a.large` (2 vCPUs, 8 GiB RAM) is sufficient. Size up based on the number of vectors you plan to store:

| Approximate vector count | Recommended flavor | Notes |
|---|---|---|
| Up to 1 million | `m2a.large` (2 vCPU, 8 GiB) | Comfortable for dev and small prod |
| 1–10 million | `m2a.xlarge` (4 vCPU, 16 GiB) | Add a block volume for storage |
| 10 million+ | `m2a.2xlarge` (8 vCPU, 32 GiB) or larger | Consider dedicated storage volume |

Vector dimensions affect memory: 1 million vectors at 768 dimensions uses about 3 GB of RAM. Use the [Qdrant capacity calculator](https://qdrant.tech/documentation/guides/capacity-planning/) for precise estimates.

## Step 1: Create a VM

<CreateVmConsole flavor="m2a.large" securityPorts="22 (SSH). Port 6333 should stay restricted; see Step 4" />

## Step 2: Install Docker

<InstallDocker />

Verify the installation:

```bash
docker --version
```

## Step 3: Run Qdrant

Start Qdrant with persistent storage:

```bash
mkdir -p ~/qdrant-data

docker run -d \
  --name qdrant \
  --restart unless-stopped \
  -p 127.0.0.1:6333:6333 \
  -p 127.0.0.1:6334:6334 \
  -v ~/qdrant-data:/qdrant/storage:z \
  qdrant/qdrant
```

The `-p 127.0.0.1:6333:6333` binding keeps the API on loopback. Data in `~/qdrant-data` persists across container restarts.

Verify Qdrant is running:

```bash
curl http://localhost:6333/
```

Expected response:

```json
{"title":"qdrant - vector search engine","version":"..."}
```

View the built-in dashboard (from your local machine via SSH tunnel; see Step 4):

```
http://localhost:6333/dashboard
```

## Step 4: Configure access

**For applications on the same VM**, use `http://localhost:6333` directly; no configuration needed.

**For applications on other VMs in the same project network**, connect using the VM's private IP. Update the Docker binding to listen on the private interface:

```bash
docker stop qdrant && docker rm qdrant

docker run -d \
  --name qdrant \
  --restart unless-stopped \
  -p PRIVATE_IP:6333:6333 \
  -p PRIVATE_IP:6334:6334 \
  -v ~/qdrant-data:/qdrant/storage:z \
  qdrant/qdrant
```

Replace `PRIVATE_IP` with the VM's private network address (visible in the console or via `ip a`). Update your project's security group to allow TCP on port 6333 from your application VMs' private CIDRs only.

**For access from your local machine**, use an SSH tunnel:

```bash
ssh -L 6333:localhost:6333 ubuntu@YOUR_FLOATING_IP -N
```

Then open `http://localhost:6333/dashboard` in your browser.



Do not expose port 6333 to `0.0.0.0/0` without authentication. Qdrant's open-source version has no built-in auth. If you need internet-facing access, place Qdrant behind a reverse proxy with authentication, or use [Qdrant's API key support](https://qdrant.tech/documentation/guides/security/) in the configuration file.



## Step 5: Create a collection and insert vectors

Test the API with a small example using Python.

Ubuntu 24.04 enforces [PEP 668](https://peps.python.org/pep-0668/), which blocks `pip install` against the system Python. Create a virtual environment first:

```bash
sudo apt-get update
sudo apt-get install -y python3-venv

python3 -m venv ~/qdrant-venv
source ~/qdrant-venv/bin/activate
```

Install the Qdrant client inside the venv:

```bash
pip install qdrant-client
```

Create a collection, insert sample vectors, and run a query. The query call returns a `QueryResponse` whose `.points` attribute is the list of matched points.

```python
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams, PointStruct

client = QdrantClient(url="http://localhost:6333")

# Create a collection with 384-dimensional vectors (all-MiniLM-L6-v2 output size)
client.create_collection(
    collection_name="documents",
    vectors_config=VectorParams(size=384, distance=Distance.COSINE),
)

# Insert sample vectors
client.upsert(
    collection_name="documents",
    points=[
        PointStruct(id=1, vector=[0.1] * 384, payload={"text": "Quake AI compute documentation"}),
        PointStruct(id=2, vector=[0.2] * 384, payload={"text": "OpenStack networking guide"}),
    ],
)

# Search
results = client.query_points(
    collection_name="documents",
    query=[0.15] * 384,
    limit=2,
).points
for r in results:
    print(r.id, r.score, r.payload)
```

For a complete RAG implementation using Ollama for embeddings, see [Build a RAG pipeline](/resources/deployments/deploy-rag-pipeline).

## Next steps

- [Run a local LLM](/resources/deployments/run-local-llm): add Ollama on the same VM or a separate instance for retrieval-augmented generation
- [Build a RAG pipeline](/resources/deployments/deploy-rag-pipeline): ingest documents, embed with Ollama, store in Qdrant, and query the result

## Clean up

Stop and remove the container, then delete the instance when finished:

```bash
docker stop qdrant && docker rm qdrant
```

## Troubleshooting

**Container exits immediately**: Run `docker logs qdrant` to see the error. Storage permission issues (`chmod 777 ~/qdrant-data` or a `:z` label on the volume mount) are the most common cause on SELinux-enabled systems.

**Out of memory on insert**: Qdrant loads vectors into RAM for indexing. If you're inserting large batches, increase the instance flavor or insert in smaller chunks.

**Dashboard is blank**: The dashboard is a single-page app served from `/dashboard`. If the page loads but shows no data, check that collections exist with `curl http://localhost:6333/collections`.
