AI & LLMs • AI Agents & RAG

RAG Tutorial in Python with Ollama (Local, Free, GitHub Code)

RAG Tutorial Python with Ollama

Ask a question about a 30-page PDF and get an answer pulled from the actual text — on your own computer, with no API key and nothing uploaded anywhere. That’s what you’ll build here, in about 90 lines of plain Python. No LangChain, no LlamaIndex: you’ll see every step of RAG, because every step is a function you wrote.

Most RAG tutorials use OpenAI and cost money per question. This one uses Ollama for both the embeddings and the answers, so it’s free, private and works offline. I also kept the parts that went wrong while I built it — they’re near the end, because that’s where you’ll learn the most.

💻 All code is on GitHub:
github.com/techprofree/rag-python-ollama-pdf
— the full app, a no-RAG comparison script, and a requirements file.

What is RAG in simple terms?

A language model only knows what it was trained on. It has never seen your PDF. RAG — Retrieval-Augmented Generation — fixes that in two moves:

  • Retrieval: search your document for the few paragraphs most relevant to the question.
  • Generation: hand those paragraphs to the model and say “answer using only this.”

The model stops guessing and starts reading. Here is the whole pipeline we’ll build:

Step What happens Tool
1. Load Read the text out of the PDF pypdf
2. Split Cut the text into overlapping chunks plain Python
3. Embed Turn each chunk into a vector of numbers Ollama + nomic-embed-text
4. Store & retrieve Save vectors, find the closest ones to a question ChromaDB
5. Generate Answer from the retrieved chunks Ollama + llama3.2

What you need

  • Ollama installed and running — if not, do the 10-minute setup first
  • Python 3.9 or newer
  • 8 GB RAM (16 GB makes the embedding step noticeably faster)
  • A text-based PDF. I’m using [NAME OF YOUR SAMPLE PDF] — download it here if you want identical results. Scanned PDFs won’t work without OCR; more on that below.

First, see the problem: ask the model without RAG

Before building anything, ask the plain model a question only your PDF can answer. This takes 30 seconds and shows exactly what RAG fixes.

🐍 Python — no_rag_baseline.py
import sys
import ollama

question = " ".join(sys.argv[1:]) or "What is this document about?"

response = ollama.generate(model="llama3.2", prompt=question)
print(response["response"])
⬛ Shell
$ python no_rag_baseline.py "[A QUESTION ONLY YOUR PDF CAN ANSWER]"

Keep that answer in mind. We’ll ask the same question again at the end.

📄 no_rag_baseline.py on GitHub

Step 1 — Install the libraries and pull two models

⬛ Shell
$ pip install ollama chromadb pypdf

pypdf reads the PDF, chromadb is a small vector database that runs inside your script, and ollama talks to the models. Now pull the models — one for embeddings, one for answers:

⬛ Shell
$ ollama pull nomic-embed-text    # ~270 MB, turns text into vectors
$ ollama pull llama3.2            # ~2 GB, writes the answers

Why a separate embedding model? Chat models are trained to write; embedding models are trained to measure meaning. nomic-embed-text is small, fast on CPU, and good enough that most local RAG projects use it.

Step 2 — Load the PDF

🐍 Python — rag_chat.py (part 1)
import sys
import ollama
import chromadb
from pypdf import PdfReader

EMBED_MODEL = "nomic-embed-text"
CHAT_MODEL = "llama3.2"
CHUNK_SIZE = 800
CHUNK_OVERLAP = 150
TOP_K = 4


def load_pdf(path: str) -> str:
    reader = PdfReader(path)
    pages = []
    for page in reader.pages:
        text = page.extract_text() or ""   # extract_text() can return None
        if text.strip():
            pages.append(text)
    full = "\n".join(pages)
    if not full.strip():
        sys.exit("No text found. This PDF is probably scanned images - run OCR first.")
    print(f"Loaded {len(reader.pages)} pages, {len(full):,} characters")
    return full

Two details that most tutorials skip and that will bite you: extract_text() returns None on blank or image-only pages, which crashes a naive text += ... loop; and a scanned PDF gives you zero characters, so we stop with a clear message instead of silently building an empty index.

Step 3 — Split into overlapping chunks

🐍 Python — rag_chat.py (part 2)
def chunk_text(text: str, size: int = CHUNK_SIZE, overlap: int = CHUNK_OVERLAP) -> list[str]:
    chunks = []
    start = 0
    while start < len(text):
        end = start + size
        chunks.append(text[start:end])
        start = end - overlap   # step back so chunks share some text
    print(f"Split into {len(chunks)} chunks")
    return chunks

Why overlap? If a sentence is cut exactly at a chunk boundary, neither chunk contains the full idea and retrieval misses it. Sharing 150 characters between neighbours fixes most of those cases. 800 characters is a good default for a 3B model; I’ll show what happened when I used 2000 in the “what went wrong” section.

Step 4 — Embed the chunks and store them in ChromaDB

🐍 Python — rag_chat.py (part 3)
def build_index(chunks: list[str], name: str = "my_pdf"):
    client = chromadb.PersistentClient(path="./chroma_db")   # saved to disk
    try:
        client.delete_collection(name)   # start clean on every run
    except Exception:
        pass
    collection = client.create_collection(name)

    batch = 32
    for i in range(0, len(chunks), batch):
        part = chunks[i:i + batch]
        result = ollama.embed(model=EMBED_MODEL, input=part)   # one call per batch
        collection.add(
            ids=[str(i + j) for j in range(len(part))],
            embeddings=result["embeddings"],
            documents=part,
        )
        print(f"  embedded {min(i + batch, len(chunks))}/{len(chunks)}", end="\r")
    print("\nIndex ready.")
    return collection

Three choices worth knowing: ollama.embed() accepts a list, so we embed 32 chunks per call instead of one at a time — several times faster on CPU. PersistentClient writes the database to a chroma_db folder, so you could skip re-indexing on the next run. And we delete the old collection first, otherwise chunks from yesterday’s PDF stay in the index and leak into today’s answers.

Step 5 — Retrieve the best chunks and generate the answer

🐍 Python — rag_chat.py (part 4)
def ask(collection, question: str) -> str:
    q = ollama.embed(model=EMBED_MODEL, input=question)["embeddings"][0]
    hits = collection.query(query_embeddings=[q], n_results=TOP_K)
    context = "\n\n---\n\n".join(hits["documents"][0])

    prompt = f"""You are answering questions about a document.
Use ONLY the context below. If the answer is not in the context, say "I can't find that in the document."

Context:
{context}

Question: {question}
Answer:"""

    answer = ""
    for chunk in ollama.generate(model=CHAT_MODEL, prompt=prompt, stream=True):
        token = chunk["response"]
        answer += token
        print(token, end="", flush=True)
    print()
    return answer

The question is embedded with the same model as the chunks (mixing embedding models is a classic silent failure), ChromaDB returns the four closest chunks, and they go into the prompt. The line telling the model to say “I can’t find that” matters: without it, a small model fills gaps with guesses and you’re back to the hallucination you saw at the start.

Step 6 — Turn it into a chat loop

🐍 Python — rag_chat.py (part 5)
def main():
    if len(sys.argv) < 2:
        sys.exit("Usage: python rag_chat.py document.pdf")

    text = load_pdf(sys.argv[1])
    chunks = chunk_text(text)
    collection = build_index(chunks)

    print("\nChat with your PDF. Type 'exit' to quit.\n")
    while True:
        question = input("You: ").strip()
        if question.lower() in {"exit", "quit"}:
            break
        if not question:
            continue
        print("AI: ", end="", flush=True)
        ask(collection, question)
        print()


if __name__ == "__main__":
    main()

Run it:

⬛ Shell
$ python rag_chat.py document.pdf

Now ask the same question from the start of the article.

📄 Full file: rag_chat.py on GitHub — all five parts together, ready to run.

📥 One-page RAG cheat sheet (PDF): the pipeline diagram, the five steps, default chunk settings and the prompt template — download it free.

What went wrong while I built this (and the fixes)

These are the real problems from my test. If your answers look off, start here.

Chunks too big — the answer ignored the document

My first version used 2000-character chunks. Four of those is 8000 characters of context, and a 3B model loses the thread; it started answering from general knowledge again. Dropping to 800 fixed it. If your answers feel vague, make chunks smaller before you try a bigger model.

[YOUR TEST: replace the above with what you actually observed, or delete this H3 if it didn’t happen.]

Retrieval returned the wrong section

Questions that use different words than the document (“cost” vs “price”, “steps” vs “procedure”) can miss. Two cheap fixes: raise TOP_K to 6, or rephrase the question using words you know are in the PDF. The proper fix is hybrid search (keyword + vector) — that’s a later tutorial.

Tables came out as word soup

pypdf reads tables row by row with no structure, so “what was the Q3 revenue” often fails. For table-heavy PDFs use pdfplumber for extraction instead; the rest of the pipeline stays the same.

Old chunks leaking into new answers

Before I added delete_collection(), running the script on a second PDF mixed both documents. If you ever see an answer from the wrong file, that’s why.

Common errors and fixes

Error: model ‘nomic-embed-text’ not found, try pulling it first

ollama._types.ResponseError: model ‘nomic-embed-text’ not found, try pulling it first

You pulled the chat model but not the embedding model (or vice-versa). Run ollama list and pull whichever is missing.

KeyError: ’embedding’ / AttributeError on ollama.embeddings

KeyError: ’embedding’

Older tutorials use ollama.embeddings(prompt=...), which returns {"embedding": [...]}. The current function is ollama.embed(input=...) and returns {"embeddings": [[...]]} — plural, and a list of lists. Update to pip install -U ollama and use the code above.

chromadb fails to install on Windows

error: Microsoft Visual C++ 14.0 or greater is required

One of ChromaDB’s dependencies needs to compile on older Python versions. Easiest fix: use Python 3.10–3.12 from python.org (not the Microsoft Store), then pip install -U pip and retry. If it still fails, install the “Desktop development with C++” workload from Visual Studio Build Tools.

ValueError: Collection my_pdf already exists

chromadb.errors.UniqueConstraintError: Collection my_pdf already exists

You’re using create_collection without deleting the old one first. Either keep the delete_collection call from Step 4, or use client.get_or_create_collection(name).

Loaded 0 characters / “No text found”

The PDF is scanned images. Run OCR first — ocrmypdf input.pdf output.pdf is the simplest tool — then use the output file.

Embedding is very slow

On CPU, expect roughly 1–3 seconds per batch of 32 chunks; a 200-page book can take a few minutes. Make sure you’re batching (one call per chunk is 10× slower), close other heavy apps, and remember the index is saved to disk — you only pay this cost once per document.

ConnectionRefusedError on localhost:11434

Ollama isn’t running. Start it from the Start menu or run ollama serve. Full details in the Ollama errors guide.

Where to take this next

  • Several PDFs at once — loop load_pdf over a folder and store the file name in each chunk’s metadata so answers can cite their source.
  • A web interface — wrap ask() in Streamlit; it’s about 20 extra lines.
  • The LangChain version — now that you know what each step does, the framework will make sense: RAG with LangChain and Ollama (coming next).
  • Make it a final-year project — add a Flask front end and SQLite to store chat history. See AI projects with source code.
💻 Code: github.com/techprofree/rag-python-ollama-pdf. If it saved you time, a star helps other people find it.

Got a PDF that this doesn’t handle well? Describe it in the comments — table-heavy, multi-column, scanned — and I’ll add the fix to this page.

Frequently Asked Questions

What is RAG in Python?

RAG (Retrieval-Augmented Generation) is a pattern where your Python code searches your own documents for relevant passages and passes them to a language model, so the model answers from your data instead of guessing. It needs a PDF/text loader, an embedding model, a vector store and a chat model — all of which run locally with Ollama and ChromaDB.

Can I build RAG locally for free?

Yes. Ollama provides both the embedding model (nomic-embed-text) and the chat model (llama3.2) free, and ChromaDB runs inside your script. Nothing leaves your computer and there are no per-query costs.

Do I need LangChain for RAG?

No. This tutorial builds RAG in about 90 lines without any framework. LangChain and LlamaIndex are useful once you need many document types, chat history or agents, but they hide what’s happening — learn the plain version first.

Which embedding model should I use with Ollama?

nomic-embed-text is the standard choice: small, fast on CPU, and accurate enough for most documents. mxbai-embed-large is a stronger alternative if you have the RAM. Always embed questions with the same model you used for the chunks.

Why does my RAG app give wrong answers?

Usually one of four things: chunks are too large for the model, the question uses different vocabulary than the document, the PDF is scanned or table-heavy so the text extraction is poor, or the prompt doesn’t tell the model to stick to the context. The “what went wrong” section above covers each fix.

Can RAG handle large PDFs?

Yes — only the few most relevant chunks are sent to the model for each question, so a 500-page document costs the same per question as a 5-page one. The only cost that scales is the one-time embedding step, which is saved to disk.