AI & LLMs • Local LLMs

Run an LLM Locally with Ollama and Python (Step-by-Step + Code)

run an llm locally with ollama and python feature image

You can run a capable AI model on your own computer, for free, with no API key and no internet connection once the model is downloaded. This guide walks through exactly how to do it with Ollama and Python: installing on Windows, Mac and Linux, running your first model, controlling it from Python, streaming responses, working with images, and fixing the errors you’re most likely to hit along the way.

💻 All the code in this guide is on GitHub:
github.com/techprofree/ollama-python-local-llm
— five runnable scripts plus a requirements file. Clone it and follow along.

Can I run LLMs locally using Ollama?

Yes. Ollama is a free, open-source tool that downloads an open model (Llama, Gemma, Qwen, Mistral and others) and runs it on your machine with one command. It also starts a local API server, which is what lets Python talk to the model. A 3-billion-parameter model such as Llama 3.2 runs on an ordinary laptop with 8 GB of RAM and no graphics card.

Why bother, when cloud APIs exist?

  • No cost — no per-token charges, no subscription
  • Privacy — your prompts and data never leave your computer
  • Works offline — once the model is downloaded
  • No rate limits — only your hardware limits you
  • Any language or framework — Ollama exposes a plain HTTP API

What is the cheapest way to run LLMs locally?

Ollama on the computer you already own, with a small model. You do not need a GPU. Here’s what actually matters:

Your RAM Model size that fits Example
8 GB Up to ~4B parameters llama3.2 (3B), gemma3:4b, qwen3:4b
16 GB Up to ~8B parameters llama3.1:8b, qwen3:8b, mistral
32 GB+ 14B and above phi4 (14B), gemma3:27b

You also need Python 3.9 or newer and roughly 2–10 GB of free disk space per model. A dedicated NVIDIA or Apple Silicon GPU makes responses several times faster, but it is optional.

Step 1 — Install Ollama on Windows (then Mac and Linux)

Windows: Go to ollama.com/download, download the Windows installer and run it. There are no options to choose. When it finishes, Ollama starts in the background and you’ll see a small llama icon in the system tray next to the clock.

macOS: Download the Mac version, open the file and drag Ollama into Applications. Launch it once.

Linux: One command:

⬛ Shell (Linux)
$ curl -fsSL https://ollama.com/install.sh | sh

On every system, confirm it worked. Open Command Prompt or PowerShell on Windows (Terminal on Mac/Linux) and run:

⬛ Shell
$ ollama --version

A version number means you’re done. If you get 'ollama' is not recognized instead, jump to the Common errors section.

Step 2 — Download and run your first model

We’ll start with Llama 3.2 (3B). It’s small enough for 8 GB machines and good enough to be useful.

⬛ Shell
$ ollama run llama3.2

The first run downloads about 2 GB, then drops you into a chat prompt in the terminal. Type a question and press Enter. Type /bye to exit.

Other models worth trying:

⬛ Shell
$ ollama run gemma3:4b      # Google, also understands images
$ ollama run qwen3:4b       # strong at reasoning for its size
$ ollama run mistral        # 7B, needs 16 GB RAM

Step 3 — Connect Ollama to Python

Install the official Python library:

⬛ Shell
$ pip install ollama

Then the simplest possible script — one prompt, one answer:

🐍 Python — 01_basic_generate.py
import ollama

response = ollama.generate(
    model="llama3.2",
    prompt="Explain what a large language model is in two sentences.",
)

print(response["response"])

Save it as 01_basic_generate.py and run python 01_basic_generate.py. The model answers from your own machine.

📄 Full file: 01_basic_generate.py on GitHub

Step 4 — Stream responses word by word

Waiting for the full answer feels slow. Pass stream=True and print each chunk as it arrives — this is how chat apps get the typing effect:

🐍 Python — 02_streaming.py
import ollama

stream = ollama.chat(
    model="llama3.2",
    messages=[{"role": "user", "content": "Write a short poem about running AI on your own computer."}],
    stream=True,
)

for chunk in stream:
    print(chunk["message"]["content"], end="", flush=True)

print()

📄 Full file: 02_streaming.py on GitHub

Step 5 — Build a chatbot with memory

The model has no memory between calls. To hold a conversation you keep a list of messages and send the whole list every time. A system message at the top sets the assistant’s behaviour.

🐍 Python — 03_chat_loop.py
import ollama

MODEL = "llama3.2"

history = [
    {"role": "system", "content": "You are a helpful, concise assistant."}
]

print(f"Chatting with {MODEL}. Type 'exit' to quit.\n")

while True:
    user_input = input("You: ").strip()
    if user_input.lower() in {"exit", "quit"}:
        break
    if not user_input:
        continue

    history.append({"role": "user", "content": user_input})

    print("AI: ", end="", flush=True)
    reply = ""
    for chunk in ollama.chat(model=MODEL, messages=history, stream=True):
        token = chunk["message"]["content"]
        reply += token
        print(token, end="", flush=True)
    print("\n")

    history.append({"role": "assistant", "content": reply})

Ask it something, then ask a follow-up like “explain that more simply” — it remembers what you were talking about.

📄 Full file: 03_chat_loop.py on GitHub

Step 6 — Describe images with a multimodal model

Some models accept images as well as text. Pull one that does, then pass the image path in the images field:

⬛ Shell
$ ollama pull gemma3:4b
🐍 Python — 04_multimodal.py
import sys
import ollama

image_path = sys.argv[1]   # run:  python 04_multimodal.py photo.jpg

response = ollama.chat(
    model="gemma3:4b",
    messages=[{
        "role": "user",
        "content": "Describe this image in detail. What objects and text can you see?",
        "images": [image_path],
    }],
)

print(response["message"]["content"])

[YOUR TEST: one sentence — which image you tried and whether the description was accurate.]

📄 Full file: 04_multimodal.py on GitHub

Step 7 — Use Ollama through the OpenAI-compatible API

Ollama also serves an OpenAI-style endpoint at http://localhost:11434/v1. Any code written for the OpenAI SDK — and most tutorials on the internet — can point at your local model by changing two lines:

⬛ Shell
$ pip install openai
🐍 Python — 05_openai_compatible.py
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",
    api_key="ollama",   # required by the SDK, ignored by Ollama
)

completion = client.chat.completions.create(
    model="llama3.2",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Give me three tips for learning Python faster."},
    ],
)

print(completion.choices[0].message.content)

This is the easiest way to run LangChain, LlamaIndex or any OpenAI-based project locally for free.

📄 Full file: 05_openai_compatible.py on GitHub

The raw REST API (any language)

Under the hood, everything above is HTTP. You can call it with requests, from JavaScript, from a shell — anything:

🐍 Python
import requests

r = requests.post(
    "http://localhost:11434/api/generate",
    json={"model": "llama3.2", "prompt": "What is Ollama?", "stream": False},
)

print(r.json()["response"])

Ollama commands you’ll use every day

⬛ Shell
$ ollama list              # installed models
$ ollama pull qwen3:4b     # download without running
$ ollama ps                # models currently loaded in memory
$ ollama rm mistral        # delete a model to free disk space
$ ollama show llama3.2     # model details (context length, parameters)
$ ollama serve             # start the server manually if it isn't running

The Python library mirrors these: ollama.list(), ollama.pull(), ollama.show(), plus ollama.embed() for embeddings when you move on to search and RAG.

Which Ollama LLM is best for coding?

For coding help, use a model trained specifically for it. The coder variants of Qwen are the usual pick on consumer hardware — qwen2.5-coder:7b fits in 16 GB and the :1.5b / :3b sizes fit in 8 GB. DeepSeek’s coder and reasoning models are the other common choice if you have more RAM. General models like Llama 3.2 can write code too, just less reliably.

Best Ollama models to run locally

Model names on ollama.com/library change often — check the exact tag before pulling. As of this writing:

Model (tag) Best for RAM needed
llama3.2 (3B) Best starting point, fast on CPU 8 GB
gemma3:4b Text + images on low-end machines 8 GB
qwen3:4b / qwen3:8b Reasoning and multilingual 8 / 16 GB
llama3.1:8b Higher quality general chat 16 GB
mistral (7B) Fast, good instruction following 16 GB
qwen2.5-coder:7b Coding and code explanation 16 GB
phi4 (14B) Strong reasoning if you have the RAM 32 GB
deepseek-r1:8b Step-by-step reasoning tasks 16 GB

Rule of thumb: parameters in billions × ~1.2 ≈ GB of RAM needed for a standard 4-bit model. If a model is bigger than your free RAM it will still run, but from disk — painfully slowly.

Common Ollama errors and fixes

These are the errors you’re most likely to see on Windows. Each one is copied exactly so you can match it.

‘ollama’ is not recognized as an internal or external command

‘ollama’ is not recognized as an internal or external command, operable program or batch file.

Command Prompt was open before Ollama was installed, so it doesn’t know the new PATH. Close the window and open a new one. If it still fails, log out and back in, or check that C:\Users\<you>\AppData\Local\Programs\Ollama is in your PATH.

ConnectionRefusedError / connection refused on localhost:11434

ConnectionError: [Errno 111] Connection refused
httpx.ConnectError: [WinError 10061] No connection could be made because the target machine actively refused it

The Ollama server isn’t running. On Windows look for the tray icon; if it’s missing, start Ollama from the Start menu. On any system you can start it manually with ollama serve in a separate terminal and leave it open.

Error: model ‘llama3.2’ not found, try pulling it first

ollama._types.ResponseError: model ‘llama3.2’ not found, try pulling it first

Your Python code asks for a model you haven’t downloaded, or the tag doesn’t match. Run ollama list to see exact names, then ollama pull llama3.2. Tags matter — gemma3 and gemma3:4b are different downloads.

ModuleNotFoundError: No module named ‘ollama’

ModuleNotFoundError: No module named ‘ollama’

The library is installed in a different Python than the one running your script — common when you have Python from the Microsoft Store and from python.org. Install it for the interpreter you’re using: python -m pip install ollama. In VS Code, check the interpreter shown in the bottom-right corner.

Ollama is very slow / not using the GPU

time=… level=INFO msg=”no compatible GPUs were discovered”

Run ollama ps while a model is loaded. The PROCESSOR column shows 100% CPU if the GPU isn’t being used. For NVIDIA cards, update to the latest driver; CUDA itself ships with Ollama. For AMD on Windows, support is limited to recent cards. If you have no GPU, the fix is a smaller model (llama3.2 rather than an 8B) and closing memory-hungry apps.

Out of memory / model keeps crashing

Error: llama runner process has terminated: … out of memory

The model doesn’t fit in RAM (or VRAM). Use the smaller tag of the same model (:4b instead of :12b), close browsers and IDEs, or shorten the context with options={"num_ctx": 2048} in your chat() call.

Antivirus or firewall blocks Ollama on Windows

Some security suites flag the first launch of ollama.exe or block port 11434. If ollama --version works but Python can’t connect, add Ollama to the allowed list in your antivirus and permit it through Windows Defender Firewall for private networks.

What to build next

  • Chat with your own PDFs — combine Ollama with embeddings and a vector database. That’s our next tutorial: Build a RAG app in Python with Ollama.
  • A private coding assistant using qwen2.5-coder and the chat loop above.
  • A final-year project — a local chatbot with a Flask front end and SQLite database (see our AI projects with source code).

If you’re new to Python itself, start with our Python beginner-to-advanced guide first — everything here assumes you can run a script and install a package.

💻 Grab the code: github.com/techprofree/ollama-python-local-llm. If it helped, a star on the repo helps other people find it.

Hit an error that isn’t listed above? Paste the exact message in the comments and I’ll add the fix to this page.

Frequently Asked Questions

Is it possible to run an LLM locally?

Yes. With Ollama you can run open models such as Llama 3.2, Gemma 3 and Qwen 3 on a normal laptop. Small models (3–4B parameters) run on 8 GB of RAM without a GPU; larger ones need 16–32 GB.

Is running an LLM locally free?

Completely. Ollama and the models in its library are free to download and use. There are no API keys, subscriptions or per-token charges, and you can send unlimited requests.

Do I need a GPU to run Ollama?

No. Ollama runs on the CPU. A GPU (NVIDIA, or Apple Silicon on Mac) makes responses several times faster, but a 3B model is usable on CPU alone.

Can I use Ollama with Python?

Yes — install the official library with pip install ollama and you can chat with a local model in five lines. Ollama also offers an OpenAI-compatible endpoint, so code written for the OpenAI SDK works by changing the base URL.

Which Ollama model should I start with?

llama3.2 for general use on 8 GB machines, gemma3:4b if you want image support, and qwen2.5-coder for programming help. Move up to 8B models once you have 16 GB of RAM.

Is a local LLM private?

Yes. The model runs entirely on your computer and nothing is sent to an external server, which makes it suitable for confidential documents and code.