# Build a RAG Pipeline with Ollama and Qdrant

Source: https://docs.quake.ai/resources/deployments/deploy-rag-pipeline
Markdown: https://docs.quake.ai/resources/deployments/deploy-rag-pipeline.md

---

# Build a RAG Pipeline with Ollama and Qdrant

Stand up a retrieval-augmented generation pipeline on Quake AI by composing [Ollama](/resources/deployments/run-local-llm) inference with [Qdrant](/resources/deployments/deploy-vector-database) vector search. You ingest documents, embed chunks with Ollama, store vectors in Qdrant, and query the stack from Python. Everything runs on your infrastructure; no external model API calls.

## Prerequisites

- Ollama running with at least one generative model (`llama3.2` or similar): [Run a local LLM](/resources/deployments/run-local-llm)
- Qdrant running at `http://localhost:6333`: [Deploy a vector database](/resources/deployments/deploy-vector-database)
- Python 3.10+ on the host

This deployment assumes Ollama and Qdrant on the same instance. For a split layout, replace `localhost` with the appropriate private IPs.

<PricingCompanion
  components={[
    { kind: "template", slug: "qdrant", required: true },
    { kind: "primitive", required: true, label: "Ollama and pipeline host", vm: { flavor: "m2a.xlarge" } },
  ]}
/>

## Architecture

<Figure size="lg" caption="RAG pipeline: the ingestion path stores document embeddings, the query path retrieves top-k context for the generative model.">

```d2
direction: down

ingestion: Ingestion path {
  docs: Documents
  split: Text splitter
  iembed: Ollama (embed)

  docs -> split -> iembed
}

query: Query path {
  user: User query
  qembed: Ollama (embed)
  generate: Ollama (generate)
  answer: Answer

  user -> qembed
  generate -> answer
}

qdrant: Qdrant {shape: cylinder}

ingestion.iembed -> qdrant: store vectors
query.qembed -> qdrant: search
qdrant -> query.generate: top-k context
```

</Figure>

The embedding model and the generative model are separate. You will pull `nomic-embed-text` for embeddings and use `llama3.2` (or your preferred model) for generation.

## Step 1: Pull the embedding model

```bash
ollama pull nomic-embed-text
```

`nomic-embed-text` produces 768-dimensional vectors and is optimized for retrieval tasks. It runs efficiently alongside a generative model on the same Ollama instance.

## Step 2: install dependencies

Ubuntu 24.04 enforces [PEP 668](https://peps.python.org/pep-0668/), which blocks `pip install` against the system Python. Create a virtual environment first:

```bash
sudo apt-get update
sudo apt-get install -y python3-venv

python3 -m venv ~/rag-venv
source ~/rag-venv/bin/activate
```

Install the langchain packages and the Qdrant client into the venv:

```bash
pip install \
  langchain-core langchain-text-splitters \
  langchain-community langchain-qdrant \
  langchain-ollama qdrant-client
```



The `langchain-community` package is [sunset upstream](https://github.com/langchain-ai/langchain-community/issues/674); the LangChain project recommends the standalone integration packages (`langchain-qdrant`, `langchain-ollama`, `langchain-text-splitters`) instead. This deployment still installs `langchain-community` because the file-based document loaders (`DirectoryLoader`, `TextLoader`) used in Step 3 currently live there. If you switch to other loaders (for example, the new `langchain-unstructured` package), you no longer need `langchain-community`.



For other document formats you may also need:

```bash
pip install pypdf unstructured          # PDF and HTML documents
pip install python-docx                  # Word documents
```

## Step 3: ingest documents

Create `ingest.py`:

```python
from langchain_community.document_loaders import (
    DirectoryLoader,
    TextLoader,
)
from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_ollama import OllamaEmbeddings
from langchain_qdrant import QdrantVectorStore
from qdrant_client import QdrantClient
from qdrant_client.models import Distance, VectorParams

OLLAMA_URL = "http://localhost:11434"
QDRANT_URL = "http://localhost:6333"
COLLECTION_NAME = "documents"
DOCS_DIR = "./docs"

# Load plain text files from ./docs/
loader = DirectoryLoader(DOCS_DIR, glob="**/*.txt", loader_cls=TextLoader)
documents = loader.load()
print(f"Loaded {len(documents)} documents")

# Split into overlapping chunks
splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,
    chunk_overlap=50,
    separators=["\n\n", "\n", ". ", " "],
)
chunks = splitter.split_documents(documents)
print(f"Split into {len(chunks)} chunks")

# Set up embedding model
embeddings = OllamaEmbeddings(model="nomic-embed-text", base_url=OLLAMA_URL)

# Create Qdrant collection
client = QdrantClient(url=QDRANT_URL)
if not client.collection_exists(COLLECTION_NAME):
    client.create_collection(
        collection_name=COLLECTION_NAME,
        vectors_config=VectorParams(size=768, distance=Distance.COSINE),
    )

# Ingest: this calls Ollama to embed each chunk
vectorstore = QdrantVectorStore.from_documents(
    chunks,
    embeddings,
    url=QDRANT_URL,
    collection_name=COLLECTION_NAME,
)
print(f"Ingested {len(chunks)} chunks into Qdrant collection '{COLLECTION_NAME}'")
```

Create a `docs/` directory with your text files, then run:

```bash
mkdir -p docs
echo "Quake AI provides compute, networking, and storage on OpenStack." > docs/overview.txt
python3 ingest.py
```

Ingestion speed depends on the number of chunks and Ollama's throughput. Embedding is faster than generation; expect about 50–200 chunks per second on `m2a.xlarge`.

## Step 4: Query the pipeline

Create `query.py`:

```python
from langchain_ollama import OllamaEmbeddings, OllamaLLM
from langchain_qdrant import QdrantVectorStore
from langchain_core.prompts import PromptTemplate
from qdrant_client import QdrantClient

OLLAMA_URL = "http://localhost:11434"
QDRANT_URL = "http://localhost:6333"
COLLECTION_NAME = "documents"
GENERATIVE_MODEL = "llama3.2"

# Connect to the existing collection
embeddings = OllamaEmbeddings(model="nomic-embed-text", base_url=OLLAMA_URL)
client = QdrantClient(url=QDRANT_URL)
vectorstore = QdrantVectorStore(
    client=client,
    collection_name=COLLECTION_NAME,
    embedding=embeddings,
)

# Configure the retriever to return the top 3 chunks
retriever = vectorstore.as_retriever(search_kwargs={"k": 3})

# Prompt template that instructs the model to use retrieved context
prompt_template = PromptTemplate(
    input_variables=["context", "question"],
    template=(
        "Answer the question using only the context provided. "
        "If the context does not contain enough information, say so.\n\n"
        "Context:\n{context}\n\n"
        "Question: {question}\n\n"
        "Answer:"
    ),
)

# Generative model
llm = OllamaLLM(model=GENERATIVE_MODEL, base_url=OLLAMA_URL)

# Run a query
question = "What does Quake AI provide?"

# Retrieve the top-k context chunks
docs = retriever.invoke(question)
context = "\n\n".join(doc.page_content for doc in docs)

# Format the prompt and ask the model
prompt = prompt_template.format(context=context, question=question)
answer = llm.invoke(prompt)

print("Answer:", answer)
print("\nSources:")
for doc in docs:
    print(" -", doc.metadata.get("source", "unknown"), doc.page_content[:100])
```

Run the query:

```bash
python3 query.py
```

The output shows the generated answer and the source chunks used to produce it.

## Step 5: Serve the pipeline as an API (optional)

For integration into applications, wrap the query logic in a FastAPI endpoint. This snippet assumes you keep the `retriever`, `prompt_template`, and `llm` setup from Step 4 in the same module.

```python
from fastapi import FastAPI
from pydantic import BaseModel

app = FastAPI()

class QueryRequest(BaseModel):
    question: str

class QueryResponse(BaseModel):
    answer: str
    sources: list[str]

@app.post("/query", response_model=QueryResponse)
async def query_endpoint(req: QueryRequest):
    docs = retriever.invoke(req.question)
    context = "\n\n".join(doc.page_content for doc in docs)
    prompt = prompt_template.format(context=context, question=req.question)
    answer = llm.invoke(prompt)
    sources = [doc.metadata.get("source", "") for doc in docs]
    return QueryResponse(answer=answer, sources=sources)
```

Install FastAPI and run:

```bash
pip install fastapi uvicorn
uvicorn app:app --host 0.0.0.0 --port 8000
```

Your RAG pipeline is now accessible at `http://YOUR_VM_IP:8000/query`. Open port 8000 in your security group (restricted to your application VMs' IPs).

## Loading PDF and other document formats

Replace the `DirectoryLoader` with the appropriate loader:

```python
from langchain_community.document_loaders import PyPDFLoader, UnstructuredHTMLLoader

# PDFs
loader = PyPDFLoader("./docs/manual.pdf")

# HTML
loader = UnstructuredHTMLLoader("./docs/page.html")

documents = loader.load()
```

LangChain supports dozens of loaders. See the [LangChain document loaders documentation](https://python.langchain.com/docs/integrations/document_loaders/) for the full list.

## Incremental ingestion

Re-running `ingest.py` adds new documents to the existing collection without replacing prior vectors, because `QdrantVectorStore.from_documents` defaults to additive writes against an existing collection. To update a document, delete the vectors associated with its source path from Qdrant first:

```python
client.delete(
    collection_name=COLLECTION_NAME,
    points_selector=Filter(
        must=[FieldCondition(key="metadata.source", match=MatchValue(value="./docs/overview.txt"))]
    ),
)
```

## Troubleshooting

**Embedding is slow**: `nomic-embed-text` is CPU-efficient, but embedding large corpora takes time on CPU. Batch size matters: LangChain's default is 1 embedding per call. For bulk ingestion, use the `OllamaEmbeddings.embed_documents()` method directly with larger batches.

**Answers reference things not in the documents**: The model may still draw on training data. Strengthen the prompt: add "Do not use any information outside the provided context." as an instruction.

**Out of memory during ingestion**: Reduce the chunk size or process documents in smaller batches by splitting the file list and calling `QdrantVectorStore.from_documents()` multiple times.

**Qdrant connection refused**: Check that the Qdrant container is running (`docker ps`) and that it is listening on port 6333 (`curl http://localhost:6333/`).

## Next steps

- [Run a local LLM with Ollama](/resources/deployments/run-local-llm)
- [Deploy a vector database with Qdrant](/resources/deployments/deploy-vector-database)
- [Add a browser interface to Ollama](/resources/deployments/add-browser-ui-to-ollama): interactive model testing without code

## Clean up

Stop the FastAPI process if you started Step 5, remove test collections from Qdrant, and delete the instance when you no longer need the pipeline.
