Qdrant vector database
Qdrant vector database
This pattern composes Compute, Network, and Block Storage into a private vector database.
What this template does#
Provisions a single instance running Qdrant in Docker on a private network, for semantic search, embeddings storage, and retrieval-augmented generation (RAG):
- Compute instance running the Qdrant container, with the API bound to the private IP only (not
0.0.0.0) - Private network, subnet, router, and security group; the REST and gRPC ports are reachable only from
private_cidr - No floating IP, so the database is not exposed to the public internet. SSH is the only rule open to the internet.
- API key authentication set from
qdrant_api_key, read from a config file mounted read-only - Vector storage on a block volume mounted at
/qdrant/storage, so data survives container and instance restarts
Set qdrant_api_key to a strong value when you apply; the template ships no default key. Once the key is set, Qdrant requires it in the api-key header on every request.
Parameters#
| Parameter | Description | Default |
|---|---|---|
key_name | SSH keypair name (must already exist) | No default |
qdrant_api_key | API key required on every request, no default | none |
qdrant_image | Engine image | qdrant/qdrant:v1.12.4 |
flavor_name | Instance size | m2a.large |
image_name | Operating system image | Ubuntu-24.04 |
app_name | Display name prefix and container name | qdrant |
http_port | REST API port | 6333 |
grpc_port | gRPC API port | 6334 |
volume_size | Vector storage volume size in GiB | 10 |
external_network | External network for the router gateway | PublicStatic |
private_cidr | Private subnet CIDR; the only range allowed to reach the API ports | 10.30.0.0/24 |
When to use this pattern#
Run a vector database on a single private instance, reached by an app on the same private network for semantic search or RAG. For a relational datastore, use self-managed PostgreSQL. For the app tier that queries this database, see the Next.js app template or Containerized app template.
Estimated cost#
Monthly cost estimate
Pricing calculator ↗Sized as a custom package on dedicated vCPU.
Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.
What each resource is for
Qdrant
m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.
Included in baseline
m2a.large
2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps
Compute + RAM rate basis
2 vCPU + 8 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.
Block storage (30 GiB)
30 GiB at $0.08/GiB/mo
Package promotional discount
Flat −$5.00/mo on the custom package (same promotion as named plans).
Included at no charge
These line items are zero on Quake AI. Many other providers meter them separately.
Data transfer (inbound and outbound)
Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.
AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.
Learn morePrivate networking
Private networks, subnets, Neutron routers, and security groups are included with the plan.
VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.
Control-plane API requests
OpenStack API calls for provisioning and management are included.
Some managed services on other clouds meter API calls or charge for premium control-plane features.
Dev/test vs production
Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.
Dev/test on shared CPU
Burstable s1a flavors; suited to prototyping and low or bursty load.
Production on dedicated CPU
The headline estimate above; predictable steady-load performance.
Saves $49.50/mo while you build on shared CPU.
Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.
Pricing data last validated: . For current rates, check quake.ai/pricing.
Template source#
Show source (7 files)Hide source
data "openstack_images_image_v2" "os" {
name = var.image_name
most_recent = true
}
data "openstack_networking_network_v2" "external" {
name = var.external_network
}
resource "openstack_networking_network_v2" "private" {
name = "${var.app_name}-private"
admin_state_up = true
}
resource "openstack_networking_subnet_v2" "private" {
name = "${var.app_name}-private-sn"
network_id = openstack_networking_network_v2.private.id
cidr = var.private_cidr
ip_version = 4
enable_dhcp = true
dns_nameservers = ["8.8.8.8", "8.8.4.4"]
}
resource "openstack_networking_router_v2" "router" {
name = "${var.app_name}-router"
external_network_id = data.openstack_networking_network_v2.external.id
admin_state_up = true
}
resource "openstack_networking_router_interface_v2" "private" {
router_id = openstack_networking_router_v2.router.id
subnet_id = openstack_networking_subnet_v2.private.id
}
resource "openstack_networking_secgroup_v2" "qdrant" {
name = "${var.app_name}-sg"
description = "SSH from any IPv4; Qdrant API ports from the private CIDR only"
}
resource "openstack_networking_secgroup_rule_v2" "http" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = var.http_port
port_range_max = var.http_port
remote_ip_prefix = var.private_cidr
security_group_id = openstack_networking_secgroup_v2.qdrant.id
}
resource "openstack_networking_secgroup_rule_v2" "grpc" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = var.grpc_port
port_range_max = var.grpc_port
remote_ip_prefix = var.private_cidr
security_group_id = openstack_networking_secgroup_v2.qdrant.id
}
resource "openstack_networking_secgroup_rule_v2" "ssh" {
direction = "ingress"
ethertype = "IPv4"
protocol = "tcp"
port_range_min = 22
port_range_max = 22
remote_ip_prefix = "0.0.0.0/0"
security_group_id = openstack_networking_secgroup_v2.qdrant.id
}
resource "openstack_blockstorage_volume_v3" "qdrant_data" {
name = "${var.app_name}-data"
size = var.volume_size
}
resource "openstack_networking_port_v2" "qdrant" {
name = "${var.app_name}-port"
network_id = openstack_networking_network_v2.private.id
security_group_ids = [openstack_networking_secgroup_v2.qdrant.id]
fixed_ip {
subnet_id = openstack_networking_subnet_v2.private.id
}
depends_on = [openstack_networking_router_interface_v2.private]
}
resource "openstack_compute_instance_v2" "qdrant" {
name = var.app_name
flavor_name = var.flavor_name
key_pair = var.key_name
user_data = templatefile("${path.module}/cloud-init/qdrant.yaml", {
qdrant_image = var.qdrant_image
qdrant_api_key = var.qdrant_api_key
http_port = var.http_port
grpc_port = var.grpc_port
app_name = var.app_name
})
block_device {
uuid = data.openstack_images_image_v2.os.id
source_type = "image"
destination_type = "volume"
volume_size = 20
boot_index = 0
delete_on_termination = true
}
network {
port = openstack_networking_port_v2.qdrant.id
}
}
resource "openstack_compute_volume_attach_v2" "qdrant_data" {
instance_id = openstack_compute_instance_v2.qdrant.id
volume_id = openstack_blockstorage_volume_v3.qdrant_data.id
}
variable "key_name" {
description = "SSH keypair name (must already exist in your project)"
type = string
}
variable "qdrant_api_key" {
description = "API key Qdrant requires on every request (service.api_key). Required, no default, so no credential ships with the template."
type = string
sensitive = true
}
variable "qdrant_image" {
description = "Container image for the Qdrant vector database engine"
type = string
default = "qdrant/qdrant:v1.12.4"
}
variable "flavor_name" {
description = "Instance size. Qdrant holds vectors in RAM for indexing, so size by vector count; m2a.large (2 vCPU, 8 GiB) suits dev and small production."
type = string
default = "m2a.large"
}
variable "image_name" {
description = "Operating system image"
type = string
default = "Ubuntu-24.04"
}
variable "app_name" {
description = "Display name prefix for compute and network resources, and the container name"
type = string
default = "qdrant"
}
variable "http_port" {
description = "REST API port, bound to the private interface only"
type = number
default = 6333
}
variable "grpc_port" {
description = "gRPC API port, bound to the private interface only"
type = number
default = 6334
}
variable "volume_size" {
description = "Block volume size in GiB for vector storage, mounted at /qdrant/storage"
type = number
default = 10
}
variable "external_network" {
description = "Shared external network for router gateway and floating IPs; defaults to PublicStatic (persisted FIP / production pattern). Override with PublicEphemeral for ephemeral demos."
type = string
default = "PublicStatic"
}
variable "private_cidr" {
description = "CIDR for the private tenant network. The API ports are reachable only from this range."
type = string
default = "10.30.0.0/24"
}
output "instance_id" {
description = "ID of the compute instance running Qdrant"
value = openstack_compute_instance_v2.qdrant.id
}
output "private_ip" {
description = "Private IP address of the Qdrant instance. Reach the API from peers on the private network."
value = openstack_compute_instance_v2.qdrant.access_ip_v4
}
output "connection_string_shape" {
description = "REST API endpoint shape for clients. Send the API key in the api-key header; the secret is not emitted here."
value = "http://${openstack_compute_instance_v2.qdrant.access_ip_v4}:${var.http_port}"
}
terraform {
required_version = ">= 1.6.0"
required_providers {
openstack = {
source = "terraform-provider-openstack/openstack"
version = "~> 2.0"
}
}
}
provider "openstack" {}
# Required: SSH keypair must already exist in your project
key_name = "YOUR_KEY_NAME"
# Required: the API key Qdrant requires on every request. No default ships with the template.
qdrant_api_key = "YOUR_STRONG_API_KEY"
# Optional: engine image and sizing
# qdrant_image = "qdrant/qdrant:v1.12.4"
# flavor_name = "m2a.large"
# image_name = "Ubuntu-24.04"
# app_name = "qdrant"
# Optional: API ports (bound to the private interface only)
# http_port = 6333
# grpc_port = 6334
# Optional: vector storage volume size in GiB
# volume_size = 10
# Optional: networking
# external_network = "PublicStatic"
# private_cidr = "10.30.0.0/24"
#cloud-config
package_update: true
packages:
- ca-certificates
- curl
write_files:
- path: /etc/qdrant/config.yaml
owner: root:root
permissions: "0640"
content: |
service:
host: 0.0.0.0
http_port: ${http_port}
grpc_port: ${grpc_port}
api_key: ${qdrant_api_key}
storage:
storage_path: /qdrant/storage
runcmd:
- |
set -e
# The data volume attaches as /dev/sdb on this platform (not /dev/vdb).
DEV=/dev/sdb
for i in $(seq 1 30); do [ -b "$DEV" ] && break; sleep 5; done
if ! blkid "$DEV" >/dev/null 2>&1; then mkfs.ext4 -F -L qdrantdata "$DEV"; fi
mkdir -p /qdrant/storage
mount "$DEV" /qdrant/storage
grep -q "$DEV" /etc/fstab || echo "$DEV /qdrant/storage ext4 defaults,nofail 0 2" >> /etc/fstab
PRIVATE_IP=$(hostname -I | awk '{print $1}')
curl -fsSL https://get.docker.com | sh
systemctl enable --now docker
docker pull ${qdrant_image}
docker run -d --name ${app_name} --restart unless-stopped \
-p $PRIVATE_IP:${http_port}:${http_port} \
-p $PRIVATE_IP:${grpc_port}:${grpc_port} \
-v /qdrant/storage:/qdrant/storage \
-v /etc/qdrant/config.yaml:/qdrant/config/config.yaml:ro \
${qdrant_image}
# Qdrant vector database
Single compute instance running [Qdrant](https://qdrant.tech) in Docker on a private network, for semantic search, embeddings storage, and retrieval-augmented generation (RAG). The API ports are reachable only from the private CIDR; the instance has no floating IP. This is the self-hosted equivalent of a hosted vector database service.
**Network class:** production — `external_network` defaults to `PublicStatic` for persisted floating IPs and multi-tier stacks; override with `PublicEphemeral` for ephemeral demos.
Vectors persist to a block volume mounted at `/qdrant/storage`, so data survives container and instance restarts.
## quake.yaml mapping
A `vector` service in a launch manifest resolves to this template. The launch handoff applies it before the runtime and passes the resolved connection details to the app as a named environment variable.
| `quake.yaml` field | Maps to |
| --- | --- |
| `services: [{ type: vector }]` | This template |
| resolved `private_ip` and `http_port` | A `QDRANT_URL` injected into the runtime's `container_env` (you supply the API key) |
## Prerequisites
- OpenTofu >= 1.6.0 or Terraform >= 1.6.0
- Quake AI account with OpenStack credentials
- An existing SSH keypair in your project (the value of `key_name` must match that keypair)
- A peer on the same private network to reach the API (an app instance, or a bastion). The database is not exposed on a floating IP.
## Usage
1. Clone or copy this template directory
2. Copy `terraform.tfvars.example` to `terraform.tfvars` and set `key_name` and `qdrant_api_key`
3. Source your OpenStack credentials: `source openrc.sh`
4. Initialize: `tofu init`
5. Preview: `tofu plan`
6. Apply: `tofu apply`
Read `private_ip` and `connection_string_shape` from the outputs to build the client URL. Send the API key in the `api-key` header on every request.
## How the instance runs the database
cloud-init installs Docker, formats and mounts a block volume at `/qdrant/storage`, then runs the Qdrant container:
- The API binds to the instance's private IP only (`-p PRIVATE_IP:6333:6333` and the gRPC port), not `0.0.0.0`.
- `service.api_key` is set from `qdrant_api_key`, read from a config file mounted read-only at `/qdrant/config/config.yaml`.
- Vector storage writes to `/qdrant/storage` on the attached volume.
## Security
- The security group allows the REST and gRPC ports only from `private_cidr`. SSH (22) is the only rule open to `0.0.0.0/0`, and the instance has no floating IP, so the API is not reachable from the public internet.
- `qdrant_api_key` is required with no default, so no credential ships with the template. The key is rendered into the instance's cloud-init data (as with any self-provisioning template); it is not committed to the repository.
- Qdrant requires the API key on every request once `service.api_key` is set. Keep the key in a secret store and inject it into clients at runtime.
## Variables
| Name | Type | Required | Default | Description |
| --- | --- | --- | --- | --- |
| `key_name` | string | yes | n/a | SSH keypair name (must already exist in your project) |
| `qdrant_api_key` | string | yes | n/a | API key Qdrant requires on every request (`service.api_key`) |
| `qdrant_image` | string | no | `qdrant/qdrant:v1.12.4` | Engine image |
| `flavor_name` | string | no | `m2a.large` | Instance size |
| `image_name` | string | no | `Ubuntu-24.04` | Operating system image |
| `app_name` | string | no | `qdrant` | Display name prefix and container name |
| `http_port` | number | no | `6333` | REST API port |
| `grpc_port` | number | no | `6334` | gRPC API port |
| `volume_size` | number | no | `10` | Vector storage volume size in GiB |
| `external_network` | string | no | `PublicStatic` | Persisted FIP / production default; override with `PublicEphemeral` for demos |
| `private_cidr` | string | no | `10.30.0.0/24` | Private subnet CIDR; the only range allowed to reach the API ports |
## Outputs
| Name | Description |
| --- | --- |
| `private_ip` | Private IP of the Qdrant instance |
| `connection_string_shape` | REST API endpoint shape (no secret emitted; send the API key in the `api-key` header) |
| `instance_id` | Compute instance ID |
## Scope
Single-instance Qdrant, not a distributed or highly available cluster. For horizontal scaling and replication, run Qdrant's distributed mode across multiple instances; this template provisions one node.
## Documentation
See also: [deploy a vector database tutorial](/resources/tutorials/deploy-vector-database)
Resources, parameters, and variables
key_namerequiredqdrant_api_keyrequiredqdrant_image="qdrant/qdrant:v1.12.4"flavor_name="m2a.large"image_name="Ubuntu-24.04"app_name="qdrant"http_port=6333grpc_port=6334volume_size=10external_network="PublicStatic"private_cidr="10.30.0.0/24"
Customize this pattern#
- Customize a template's image and flavor
- Add a block volume to a template
- Parameterize a template with a tfvars file
See also#
Usage Guidelines
The sample code, software libraries, command line tools, proofs of concept, templates, and other related technology on this page (including any of the foregoing that is provided by Quake AI personnel) is provided to you as Quake AI Content under the Quake AI Customer Agreement, or the relevant written agreement between you and Quake AI (whichever applies). Do not use this Quake AI Content in your production accounts, or on production or other critical data. You are responsible for testing, securing, and optimizing the Quake AI Content (such as sample code) as appropriate for production grade use based on your specific quality control practices and standards. Deploying Quake AI Content may incur Quake AI charges for creating or using Quake AI chargeable resources, such as running Compute instances or storing data in Object Storage. Your use is also subject to the Acceptable Use Policy.
For the full policy, see Usage Guidelines.
See Also
Terraform and OpenTofu on Quake AI
Prerequisite
Networks
Prerequisite
Authoring IaC templates for Quake AI
Shares: Volumes, Security Groups
Deploy an API gateway with the api-gateway template
Shares: Volumes, Security Groups
Deploy the development environment template with OpenTofu
Shares: Volumes, Security Groups