Skip to content

Build a RAG Pipeline with Ollama and Qdrant

Deployment · Updated May 2026

Build a RAG Pipeline with Ollama and Qdrant

Stand up a retrieval-augmented generation pipeline on Quake AI by composing Ollama inference with Qdrant vector search. You ingest documents, embed chunks with Ollama, store vectors in Qdrant, and query the stack from Python. Everything runs on your infrastructure; no external model API calls.

Prerequisites#

This deployment assumes Ollama and Qdrant on the same instance. For a split layout, replace localhost with the appropriate private IPs.

Monthly cost estimate

Pricing calculator ↗

Sized as a custom package on dedicated vCPU.

Starting template$195.40/mo

Monthly total for the required template above. Use the configurator below to add optional pieces and see the total update.

What each resource is for

Qdrant

m2a.large · 2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00/mo

m2a.xlarge

m2a.xlarge · 4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00/mo

Compute shown per role at custom-package rates ($29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM). The headline above is the billed total: the cheaper of a named plan and the custom package, plus add-ons.

Included in baseline

m2a.large

2 dedicated vCPU, 8 GiB RAM, 0.5 Gbps

$66.00

m2a.xlarge

4 dedicated vCPU, 16 GiB RAM, 1 Gbps

$132.00

Compute + RAM rate basis

6 vCPU + 24 GiB RAM at $29/dedicated vCPU, $7.25/shared vCPU, $1/GiB RAM (regular). Totals apply the flat −$5/mo package promotion.

—

Block storage (30 GiB)

30 GiB at $0.08/GiB/mo

$2.40

Package promotional discount

Flat −$5.00/mo on the custom package (same promotion as named plans).

$-5.00

Included at no charge

These line items are zero on Quake AI. Many other providers meter them separately.

Data transfer (inbound and outbound)

Unlimited data transfer on every plan; Quake AI does not meter per-GB egress.

AWS, GCP, and Azure meter outbound transfer per GB. DigitalOcean and Hetzner include an allowance on compute plans, then charge overage.

Learn more
$0.00

Private networking

Private networks, subnets, Neutron routers, and security groups are included with the plan.

VPC objects are usually free to create elsewhere, but NAT gateways bill hourly plus per-GB processed. Quake AI uses router SNAT with no separate NAT line item.

$0.00

Control-plane API requests

OpenStack API calls for provisioning and management are included.

Some managed services on other clouds meter API calls or charge for premium control-plane features.

$0.00

Dev/test vs production

Start on shared CPU for dev/test, then promote to dedicated for production with a flavor resize. The network, storage, and template stay the same.

Dev/test on shared CPU

Burstable s1a flavors; suited to prototyping and low or bursty load.

$46.90/mo

Production on dedicated CPU

The headline estimate above; predictable steady-load performance.

$195.40/mo

Saves $148.50/mo while you build on shared CPU.

Shared flavors carry less RAM (m2a.large (8 GiB RAM) -> s1a.small (2 GiB RAM); m2a.xlarge (16 GiB RAM) -> s1a.medium (4 GiB RAM)). A resize reboots the instance; data on attached volumes persists. Size the dedicated flavor for the RAM your production workload needs.

Pricing data last validated: . For current rates, check quake.ai/pricing.

Architecture#

Ingestion pathQuery pathQdrantDocumentsText splitterOllama (embed)User queryOllama (embed)Ollama (generate)Answer store vectorssearchtop-k context
Click to zoom
RAG pipeline: the ingestion path stores document embeddings, the query path retrieves top-k context for the generative model.

The embedding model and the generative model are separate. You will pull nomic-embed-text for embeddings and use llama3.2 (or your preferred model) for generation.

Step 1: Pull the embedding model#

bash
ollama pull nomic-embed-text

nomic-embed-text produces 768-dimensional vectors and is optimized for retrieval tasks. It runs efficiently alongside a generative model on the same Ollama instance.

Step 2: install dependencies#

Ubuntu 24.04 enforces PEP 668, which blocks pip install against the system Python. Create a virtual environment first:

bash
sudo apt-get update
sudo apt-get install -y python3-venv

python3 -m venv ~/rag-venv
source ~/rag-venv/bin/activate

Install the langchain packages and the Qdrant client into the venv:

bash
pip install \
  langchain-core langchain-text-splitters \
  langchain-community langchain-qdrant \
  langchain-ollama qdrant-client

For other document formats you may also need:

bash
pip install pypdf unstructured          # PDF and HTML documents
pip install python-docx                  # Word documents

Step 3: ingest documents#

Create ingest.py:

Python
from langchain_community.document_loaders import (
    DirectoryLoader,
    TextLoader,
)
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_ollama import OllamaEmbeddings
from langchain_qdrant import QdrantVectorStore
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams

OLLAMA_URL = "http://localhost:11434"
QDRANT_URL = "http://localhost:6333"
COLLECTION_NAME = "documents"
DOCS_DIR = "./docs"

# Load plain text files from ./docs/
loader = DirectoryLoader(DOCS_DIR, glob="**/*.txt", loader_cls=TextLoader)
documents = loader.load()
print(f"Loaded {len(documents)} documents")

# Split into overlapping chunks
splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,
    chunk_overlap=50,
    separators=["\n\n", "\n", ". ", " "],
)
chunks = splitter.split_documents(documents)
print(f"Split into {len(chunks)} chunks")

# Set up embedding model
embeddings = OllamaEmbeddings(model="nomic-embed-text", base_url=OLLAMA_URL)

# Create Qdrant collection
client = QdrantClient(url=QDRANT_URL)
if not client.collection_exists(COLLECTION_NAME):
    client.create_collection(
        collection_name=COLLECTION_NAME,
        vectors_config=VectorParams(size=768, distance=Distance.COSINE),
    )

# Ingest: this calls Ollama to embed each chunk
vectorstore = QdrantVectorStore.from_documents(
    chunks,
    embeddings,
    url=QDRANT_URL,
    collection_name=COLLECTION_NAME,
)
print(f"Ingested {len(chunks)} chunks into Qdrant collection '{COLLECTION_NAME}'")

Create a docs/ directory with your text files, then run:

bash
mkdir -p docs
echo "Quake AI provides compute, networking, and storage on OpenStack." > docs/overview.txt
python3 ingest.py

Ingestion speed depends on the number of chunks and Ollama's throughput. Embedding is faster than generation; expect about 50–200 chunks per second on m2a.xlarge.

Step 4: Query the pipeline#

Create query.py:

Python
from langchain_ollama import OllamaEmbeddings, OllamaLLM
from langchain_qdrant import QdrantVectorStore
from langchain_core.prompts import PromptTemplate
from qdrant_client import QdrantClient

OLLAMA_URL = "http://localhost:11434"
QDRANT_URL = "http://localhost:6333"
COLLECTION_NAME = "documents"
GENERATIVE_MODEL = "llama3.2"

# Connect to the existing collection
embeddings = OllamaEmbeddings(model="nomic-embed-text", base_url=OLLAMA_URL)
client = QdrantClient(url=QDRANT_URL)
vectorstore = QdrantVectorStore(
    client=client,
    collection_name=COLLECTION_NAME,
    embedding=embeddings,
)

# Configure the retriever to return the top 3 chunks
retriever = vectorstore.as_retriever(search_kwargs={"k": 3})

# Prompt template that instructs the model to use retrieved context
prompt_template = PromptTemplate(
    input_variables=["context", "question"],
    template=(
        "Answer the question using only the context provided. "
        "If the context does not contain enough information, say so.\n\n"
        "Context:\n{context}\n\n"
        "Question: {question}\n\n"
        "Answer:"
    ),
)

# Generative model
llm = OllamaLLM(model=GENERATIVE_MODEL, base_url=OLLAMA_URL)

# Run a query
question = "What does Quake AI provide?"

# Retrieve the top-k context chunks
docs = retriever.invoke(question)
context = "\n\n".join(doc.page_content for doc in docs)

# Format the prompt and ask the model
prompt = prompt_template.format(context=context, question=question)
answer = llm.invoke(prompt)

print("Answer:", answer)
print("\nSources:")
for doc in docs:
    print(" -", doc.metadata.get("source", "unknown"), doc.page_content[:100])

Run the query:

bash
python3 query.py

The output shows the generated answer and the source chunks used to produce it.

Step 5: Serve the pipeline as an API (optional)#

For integration into applications, wrap the query logic in a FastAPI endpoint. This snippet assumes you keep the retriever, prompt_template, and llm setup from Step 4 in the same module.

Python
from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

class QueryRequest(BaseModel):
    question: str

class QueryResponse(BaseModel):
    answer: str
    sources: list[str]

@app.post("/query", response_model=QueryResponse)
async def query_endpoint(req: QueryRequest):
    docs = retriever.invoke(req.question)
    context = "\n\n".join(doc.page_content for doc in docs)
    prompt = prompt_template.format(context=context, question=req.question)
    answer = llm.invoke(prompt)
    sources = [doc.metadata.get("source", "") for doc in docs]
    return QueryResponse(answer=answer, sources=sources)

Install FastAPI and run:

bash
pip install fastapi uvicorn
uvicorn app:app --host 0.0.0.0 --port 8000

Your RAG pipeline is now accessible at http://YOUR_VM_IP:8000/query. Open port 8000 in your security group (restricted to your application VMs' IPs).

Loading PDF and other document formats#

Replace the DirectoryLoader with the appropriate loader:

Python
from langchain_community.document_loaders import PyPDFLoader, UnstructuredHTMLLoader

# PDFs
loader = PyPDFLoader("./docs/manual.pdf")

# HTML
loader = UnstructuredHTMLLoader("./docs/page.html")

documents = loader.load()

LangChain supports dozens of loaders. See the LangChain document loaders documentation for the full list.

Incremental ingestion#

Re-running ingest.py adds new documents to the existing collection without replacing prior vectors, because QdrantVectorStore.from_documents defaults to additive writes against an existing collection. To update a document, delete the vectors associated with its source path from Qdrant first:

Python
client.delete(
    collection_name=COLLECTION_NAME,
    points_selector=Filter(
        must=[FieldCondition(key="metadata.source", match=MatchValue(value="./docs/overview.txt"))]
    ),
)

Troubleshooting#

Embedding is slow: nomic-embed-text is CPU-efficient, but embedding large corpora takes time on CPU. Batch size matters: LangChain's default is 1 embedding per call. For bulk ingestion, use the OllamaEmbeddings.embed_documents() method directly with larger batches.

Answers reference things not in the documents: The model may still draw on training data. Strengthen the prompt: add "Do not use any information outside the provided context." as an instruction.

Out of memory during ingestion: Reduce the chunk size or process documents in smaller batches by splitting the file list and calling QdrantVectorStore.from_documents() multiple times.

Qdrant connection refused: Check that the Qdrant container is running (docker ps) and that it is listening on port 6333 (curl http://localhost:6333/).

Next steps#

Clean up#

Stop the FastAPI process if you started Step 5, remove test collections from Qdrant, and delete the instance when you no longer need the pipeline.

Was this page helpful?