Build an AI Agent in Pure Python with Ollama (From Scratch, With Real Tool Calling)
A chatbot talks. An agent does things: it decides it needs a calculator, calls one, reads the result, and only then answers. In this guide you’ll build that in about 70 lines of plain Python on a free local model — no LangChain, no API key — using Ollama’s native function calling, so the model returns structured tool calls instead of text you have to parse with regex.
The agent gets three real tools: a calculator that can’t be hijacked, a file reader locked to one folder, and the PDF search from the RAG tutorial turned into a tool. Then we do what most tutorials skip: test which small models actually call tools correctly, and look at the ways the agent fails.
github.com/techprofree/ai-agent-python-ollama
— the agent, the tools, and a script that scores models on tool calling.
[VIDEO: embed your 2–3 minute YouTube screen recording of the agent answering three questions]
How do you build an AI agent with Ollama?
Four pieces: (1) some Python functions, (2) a JSON description of each so the model knows they exist, (3) a loop that sends the conversation plus those descriptions to ollama.chat() and runs whatever the model asks for, and (4) guard-rails so the loop can’t run forever. That’s the entire architecture — frameworks add convenience on top, not new ideas.
The loop looks like this:
| Step | Who | What happens |
|---|---|---|
| 1 | You | Send the question + tool descriptions to the model |
| 2 | Model | Either answers in text, or returns tool_calls: a tool name and arguments |
| 3 | You | Run the function, append the result as a tool message |
| 4 | — | Back to step 1 — until the model answers in text, or you hit the step limit |
What you need
- Ollama running — setup guide if you’re new
- A model that supports tool calling. On 8 GB RAM:
llama3.2orqwen3:4b. On 16 GB:llama3.1:8borqwen3:8b. - Python 3.9+, and for the PDF tool,
nomic-embed-textpluschromadbandpypdffrom the RAG post
$ ollama pull llama3.2
$ ollama pull nomic-embed-text
$ pip install ollama chromadb pypdf
Step 1 — Write the tools (plain Python functions)
Every tool takes strings and returns a string. Three tools, each chosen to teach something:
A calculator that isn’t a security hole
Most tutorials write return str(eval(expression)). Never do that: the model (or whoever is typing to it) can pass __import__('os').system('rm -rf ~') and eval will happily run it. Instead we parse the expression into a syntax tree and only allow numbers and arithmetic operators:
import ast
import operator
_OPS = {
ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,
ast.Div: operator.truediv, ast.Pow: operator.pow, ast.Mod: operator.mod,
ast.USub: operator.neg,
}
def _eval_node(node):
if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):
return node.value
if isinstance(node, ast.BinOp) and type(node.op) in _OPS:
return _OPS[type(node.op)](_eval_node(node.left), _eval_node(node.right))
if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:
return _OPS[type(node.op)](_eval_node(node.operand))
raise ValueError("only numbers and + - * / ** % are allowed")
def calculate(expression: str) -> str:
try:
tree = ast.parse(expression, mode="eval")
return str(_eval_node(tree.body))
except Exception as e:
return f"Error: {e}"
Try calculate("__import__('os')") — it returns an error string instead of running anything. Also notice every tool returns errors as text rather than raising: the model can read “Error: …” and try again, but an exception would crash the loop.
A file reader locked to one folder
from pathlib import Path
DOCS_DIR = Path(__file__).parent / "docs"
def read_file(filename: str) -> str:
path = (DOCS_DIR / Path(filename).name).resolve() # .name strips ../ tricks
if not path.is_file():
available = ", ".join(p.name for p in DOCS_DIR.iterdir()) or "none"
return f"Error: '{filename}' not found. Files available: {available}"
return path.read_text(encoding="utf-8", errors="ignore")[:4000]
Path(filename).name throws away any directory part, so ../../secrets.txt becomes secrets.txt inside docs/. And when the file isn’t found, the error lists what is there — small models recover from that far better than from a bare “not found”.
PDF search — the RAG app as a tool
This is the same load → chunk → embed → query pipeline from the RAG tutorial, wrapped in one function. It indexes the first PDF in docs/ the first time it’s called:
def search_pdf(query: str) -> str:
import ollama
col = _get_collection() # builds the ChromaDB index on first call
if col is None:
return "Error: no PDF found in docs/ folder"
q = ollama.embed(model="nomic-embed-text", input=query)["embeddings"][0]
hits = col.query(query_embeddings=[q], n_results=3)
return "\n\n---\n\n".join(hits["documents"][0])
📄 Full file with _get_collection(): tools.py on GitHub
Step 2 — Describe the tools so the model can see them
The model never sees your Python. It sees a JSON schema per tool: name, description, and parameters. The description is the single most important string in the whole project — vague descriptions are the #1 cause of the model picking the wrong tool.
TOOL_FUNCTIONS = {"calculate": calculate, "read_file": read_file, "search_pdf": search_pdf}
TOOL_SCHEMAS = [
{
"type": "function",
"function": {
"name": "calculate",
"description": "Evaluate a math expression exactly. Use for ANY arithmetic, even simple.",
"parameters": {
"type": "object",
"properties": {"expression": {"type": "string", "description": "e.g. '847 * 293'"}},
"required": ["expression"],
},
},
},
# read_file and search_pdf follow the same shape - see tools.py
]
“Use for ANY arithmetic, even simple” is there on purpose: without it, small models do easy sums in their head, get them wrong, and never touch the tool.
How do you create an AI agent in Python from scratch?
You write a loop. Here’s the whole thing — this is agent.py:
import ollama
from tools import TOOL_FUNCTIONS, TOOL_SCHEMAS
MODEL = "llama3.2"
MAX_STEPS = 6 # guard-rail: stop runaway loops
SYSTEM = """You are a helpful assistant with tools.
Use the calculate tool for any arithmetic - never do math in your head.
Use search_pdf for questions about the document, and read_file for text files.
When you have what you need, answer the user in plain text."""
def run_agent(question: str, verbose: bool = True) -> str:
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": question},
]
for step in range(1, MAX_STEPS + 1):
response = ollama.chat(model=MODEL, messages=messages, tools=TOOL_SCHEMAS)
msg = response.message
if not msg.tool_calls: # plain text = final answer
return msg.content
messages.append(msg) # keep the tool request in history
for call in msg.tool_calls:
name = call.function.name
args = call.function.arguments or {}
if name not in TOOL_FUNCTIONS: # model invented a tool
result = f"Error: unknown tool '{name}'. Available: {list(TOOL_FUNCTIONS)}"
else:
try:
result = TOOL_FUNCTIONS[name](**args)
except TypeError as e: # wrong / missing arguments
result = f"Error: bad arguments for {name}: {e}"
if verbose:
print(f" [step {step}] {name}({args}) -> {str(result)[:80]!r}")
messages.append({"role": "tool", "content": str(result), "tool_name": name})
return "Stopped: reached the step limit without a final answer."
Three things in here that the simple version in most tutorials lacks, and each one came from a real failure:
- It’s a loop, not a one-shot. Ask → tool → ask → tool → answer. Multi-step questions (“read notes.txt and multiply the number in it by 3”) need two tool calls in sequence.
- Unknown tools don’t crash it. Small models sometimes call
get_weatherorsearch_webbecause those names are common in their training data. We hand the error back and the model usually self-corrects. - MAX_STEPS. Without it, a model that keeps calling the same tool loops forever. Six is plenty for anything a 3B model can do.
Step 3 — Run it
$ python agent.py "What is 847 times 293?"
[step 1] calculate({'expression': '847 * 293'}) -> '248171'
847 times 293 is 248,171.
$ python agent.py "Read notes.txt and summarise it"
$ python agent.py "According to the PDF, what is the main topic?"
$ python agent.py "Who wrote Romeo and Juliet?" # no tool needed
Run python agent.py with no arguments for an interactive chat. The [step N] lines are the agent thinking out loud — leave them on while you’re learning; it’s the fastest way to see why an answer went wrong.
Which small models can actually call tools?
“Supports tool calling” on a model card doesn’t mean it calls the right tool with the right arguments. I ran the same five tasks against each model on an 8 GB laptop with no GPU, using compare_models.py from the repo:
| Model | T1 calc | T2 calc | T3 read_file | T4 search_pdf | T5 no tool | Score | Avg time |
|---|---|---|---|---|---|---|---|
llama3.2 |
[?] | [?] | [?] | [?] | [?] | [?]/5 | [?]s |
qwen3:4b |
[?] | [?] | [?] | [?] | [?] | [?]/5 | [?]s |
gemma3:4b |
[?] | [?] | [?] | [?] | [?] | [?]/5 | [?]s |
Run it yourself — results change with model versions and hardware, and the script prints a Markdown table you can paste anywhere:
$ ollama pull qwen3:4b
$ ollama pull gemma3:4b
$ python compare_models.py
How the agent fails (and the fix for each)
Watch the [step N] output for a while and you’ll see all of these. Each guard-rail in agent.py exists because of one.
It does the math in its head and gets it wrong
The most common failure on 3B models: “What is 847 × 293?” → “248,071” with no tool call. Fix: the system prompt line never do math in your head plus “Use for ANY arithmetic” in the tool description. Both together got llama3.2 to use the calculator reliably; either alone didn’t.
[YOUR TEST: confirm or replace with what you saw.]
It invents a tool
Models have seen thousands of weather-tool examples in training. Returning an error that lists the real tools lets the model recover on the next step. Without the if name not in TOOL_FUNCTIONS check this is a KeyError crash.
It calls the same tool forever
Usually after a tool returns an error the model doesn’t understand. MAX_STEPS ends it with a clear message instead of hanging. If you see it often, make the tool’s error messages more helpful — they’re prompts too.
It uses a tool when it shouldn’t
“Who wrote Romeo and Juliet?” → search_pdf("Romeo and Juliet"). Over-eager tool use is harmless but slow. Tightening the search_pdf description to “Use when asked about the document’s contents” cut this down.
Wrong argument names
The model called read_file(file=...) instead of filename=.... We catch the TypeError and return it as text; the model normally fixes it on the next step. Giving an example value in the parameter description (“e.g. ‘notes.txt'”) reduces this a lot.
Common errors and fixes
Error: llama3.2 does not support tools
The model you pulled has no tool-calling template. Switch to one that does — llama3.2, llama3.1, qwen3, mistral — and make sure you’re on a recent Ollama (ollama --version).
AttributeError: ‘dict’ object has no attribute ‘message’
Your ollama Python package is old and returns dicts. pip install -U ollama. (Older code that uses response["message"]["tool_calls"] still works on new versions, but the attribute style is cleaner.)
tool_calls is always empty
Three causes, in order of likelihood: the tools= argument isn’t being passed to ollama.chat(); the tool descriptions are too vague for the model to see a match; or the question doesn’t need a tool. Test with a blunt one: “Use the calculator to compute 2+2”.
ModuleNotFoundError: No module named ‘tools’
Run agent.py from inside the repo folder so Python finds tools.py next to it — cd ai-agent-python-ollama first.
search_pdf returns “no PDF found”
Put a text-based PDF in the docs/ folder. The tool indexes the first .pdf it finds there.
Very slow first search_pdf call
That call builds the whole index. It’s a one-time cost per run; see the RAG errors section for speeding up embedding.
Where to take this next
- Add a tool that writes —
save_note(text)appending to a file. Watch how the model’s behaviour changes when it can have side effects, and why you’d add a confirmation step. - Give it memory — keep
messagesbetween questions so it can refer back. - Swap in LangGraph — now the framework’s “nodes and edges” will map onto the loop you just wrote: LangGraph tutorial with Ollama (coming next).
- Final-year project version — Flask UI + SQLite log of every tool call. See AI projects with source code.
eval model output, and add a confirmation step for anything that sends, deletes or pays.compare_models.py on a different model? Post your table in the comments and I’ll add it to the page.
Frequently Asked Questions
A program where a language model can call your functions. You describe the functions, the model decides when to use one and with what arguments, your code runs it and sends the result back, and the model uses that to answer. The loop of decide → act → observe is what makes it an agent rather than a chatbot.
Yes — this guide does it in about 70 lines with only the ollama package. Frameworks are useful once you need many tools, persistence or multi-agent setups, but the core loop is simple enough to own yourself.
Llama 3.1 and 3.2, Qwen 2.5 and 3, Mistral, and most newer instruction-tuned models. Support is only half the story: the comparison table above shows how reliably each small model actually picks the right tool.
Yes. OpenAI calls it function calling, Ollama and Anthropic call it tool use; the mechanism — JSON schemas in, structured calls out — is the same.
Yes with a 3–4B model such as llama3.2 or qwen3:4b. Expect a few seconds per step on CPU. The PDF-search tool adds the embedding model (~270 MB).
Only with limits. This guide’s calculator parses expressions instead of using eval, and the file reader can’t leave its folder. Anything that writes, sends or deletes should ask you before it runs.



