Ollama is a free, open-source tool for running LLMs (large language models) on your own computer. It installs on Windows, macOS, and Linux, and a single line, ollama run model-name, takes you from downloading a model to chatting with it. While it runs, an API server listens on localhost:11434, so your own apps and other tools can call it too.

This guide sticks to how to use Ollama itself: the install command for each OS, the full command list, Modelfiles, environment variables, and minimal API examples, all checked against the official GitHub repo and the official docs (as of September 29, 2026). For local LLMs in general, see how to run a local LLM; for which model to pick, our best local LLM models comparison; and for wiring a model into your editor, the Ollama + Cline article.

LOCAL LLM RUNTIME

One command, a local LLM

— It handles almost all the setup hassle for you

$ ollama pull qwen3
$ ollama run qwen3
>>> Hi! What can you do?

✅ Free / OSS

🖥️ Win/Mac/Linux

🔌 Local API

⏱️ Minutes to set up

1. What is Ollama? The go-to local-LLM runtime

Ollama is a free, open-source tool for running local LLMs easily on your own PC. It handles the hassle—downloading models, dealing with quantization formats, configuring GPU use—behind the scenes, so all you do is "name a model and run it."

💡 In a nutshell: Ollama is "Docker for LLMs." Fetch a model with ollama pull, chat with ollama run. It also spins up a local API server, so your own apps and chat UIs can call it too.

A similar tool is LM Studio. Roughly: Ollama = CLI-first, for developers, APIs, and automation; LM Studio = GUI-first, for non-engineers getting started. Both are free and install in minutes. This article centers on Ollama (which also covers APIs and embedding); if you want a GUI, jump to Section 5.

2. How to install (Windows / macOS / Linux)

The official download page and the GitHub README offer two routes for each OS: a one-line command you paste into a terminal, or a manual installer. Either one gets you to the same place.

Windows

Open PowerShell and run this line (exactly as the official page gives it):

irm https://ollama.com/install.ps1 | iex

If you'd rather not use a command, click "Download for Windows" on the download page and run OllamaSetup.exe. Per the official Windows page, you need Windows 10 22H2 or newer (Home or Pro). No administrator rights are required; by default it installs into your home folder. The binaries alone need at least 4 GB of free space, and models can add tens to hundreds of GB. After installing, Ollama runs in the background and the ollama command works in both cmd and PowerShell.

macOS

Run this line in Terminal:

curl -fsSL https://ollama.com/install.sh | sh

To install by hand, open Ollama.dmg and drag the app into your Applications folder. On first launch, if the ollama command isn't on your PATH, the app asks for permission to create a link in /usr/local/bin. You need macOS 14 Sonoma or later; Apple M-series Macs use both CPU and GPU, while Intel (x86) Macs run on the CPU only (official macOS page).

Linux

One line installs it. To update later, run the same command again.

curl -fsSL https://ollama.com/install.sh | sh

Manual install steps without the script, and the extra package for AMD GPUs, are on the official Linux page.

Docker

The official image ollama/ollama is on Docker Hub. Here is the minimal CPU-only setup (for NVIDIA GPUs, install the NVIDIA Container Toolkit and add --gpus=all; details on the official Docker page):

docker run -d -v ollama:/root/.ollama -p 11434:11434 --name ollama ollama/ollama
docker exec -it ollama ollama run gemma3:4b

🔌 Check it works: after installing, open a new terminal and type ollama -v (or ollama --version). If it prints a version, you're set. Updates download automatically on Windows and macOS; click the taskbar or menu bar icon and choose "Restart to update" to apply them. On Linux, re-run the install line (official FAQ).

3. Ollama command list

This table covers the commands listed by the current ollama --help, matched against the official CLI reference and the definitions in the source code. Day to day, you'll mostly use the first five.

Command What it does Example
ollama run MODEL [PROMPT] Starts a model and chats with it; downloads it first if needed; with a PROMPT, it answers once and exits ollama run qwen3
ollama pull MODEL Downloads a model without starting it ollama pull gemma3:4b
ollama list Lists downloaded models and their sizes (alias: ls) ollama ls
ollama ps Lists models currently loaded in memory; the PROCESSOR column shows GPU or CPU ollama ps
ollama stop MODEL Unloads a running model from memory ollama stop qwen3
ollama rm MODEL Deletes a model to free disk space (accepts several) ollama rm gemma3:4b
ollama show MODEL Shows model details; --modelfile and --parameters reveal its settings ollama show --modelfile qwen3
ollama create MODEL Builds your own model from a Modelfile ollama create my-assistant -f Modelfile
ollama cp SOURCE DESTINATION Copies a model under a new name ollama cp qwen3 my-qwen
ollama push MODEL Uploads your model to a registry (ollama.com) ollama push your-name/my-assistant
ollama serve Starts the API server (alias: start); on Windows and macOS the background app does this for you ollama serve
ollama signin / signout Signs in to or out of ollama.com (needed for cloud models) ollama signin
ollama launch [INTEGRATION] Launches an integration such as Claude Code or Codex using Ollama models ollama launch claude

Inside a chat (after the >>> prompt), commands starting with / are available. /bye exits, /? lists them, /clear wipes the conversation context, and /set parameter num_ctx 8192 changes the context length on the spot. To type several lines at once, wrap them in """.

4. Getting and choosing models

Specify a model by name + size tag. If you omit the tag, the default latest is used. For example, llama3.2 is the same as llama3.2:latest, which is the 3B build llama3.2:3b (the official tag list shows they are the same version). If you want the smaller 1B build, ask for llama3.2:1b explicitly. Which size latest points to differs from model to model, so check the tag list before you pull. The rule of thumb: pick a size that fits in your VRAM.

# Try a lightweight model (starter, about 3.3 GB)
ollama run gemma3:4b
# Solid all-rounder (latest = 8B, about 5.2 GB)
ollama run qwen3
# For coding (latest = 30B, about 19 GB)
ollama run qwen3-coder

Sizes are as shown for each tag in the official library (as of September 29, 2026). With some models, such as qwen3-coder, leaving off the tag pulls a large build, so check the tag list first on a machine with little disk or memory. Once a model is installed, ollama show qwen3 shows its parameter count, quantization, context length, and more.

💡 Which model? Decide by use case (general / coding / your language) and size. For picks by lineage and use case, see our best local LLM models comparison; for the VRAM each size needs, see the hardware requirements article. When unsure, start small (7B class).

5. Using a GUI (Open WebUI and more)

Not a fan of the terminal? No problem—you can put a chat screen (GUI) on top of Ollama.

Open WebUI

A popular ChatGPT-style screen you connect to your local Ollama. Supports chat history, model switching, and multiple users.

Want a GUI from the start? LM Studio

A single app that handles model search, download, and chat. Ideal for non-engineers getting started. On Apple Silicon it can be fast via the MLX format.

6. Using the API (embed it in apps)

Ollama's real strength is its local API. By default the server listens on 127.0.0.1:11434, and by sending requests there, your own apps, scripts, and tools can use a local LLM. There are two main entry points.

Native API

POST localhost:11434
 /api/generate
 /api/chat

Ollama's own simple format, which also takes Ollama-specific options such as keep_alive.

OpenAI-compatible API

POST localhost:11434
 /v1/chat/completions
 /v1/responses

Reuse existing OpenAI code by just changing the endpoint.

/api/generate (a single answer)

curl http://localhost:11434/api/generate -d '{
  "model": "qwen3",
  "prompt": "Why is the sky blue?",
  "stream": false
}'

Drop "stream": false and the answer comes back bit by bit as lines of JSON (streaming is the default).

/api/chat (pass the conversation history)

curl http://localhost:11434/api/chat -d '{
  "model": "qwen3",
  "messages": [{
    "role": "user",
    "content": "Why is the sky blue?"
  }],
  "stream": false
}'

Those two are written for macOS and Linux shells. Windows PowerShell handles quotes differently, so use this form from the official Windows page:

(Invoke-WebRequest -method POST -Body '{"model":"qwen3", "prompt":"Why is the sky blue?", "stream": false}' -uri http://localhost:11434/api/generate ).Content | ConvertFrom-json

OpenAI-compatible (/v1)

With the official OpenAI library, pointing base_url at Ollama is all it takes. According to the official docs, the API key is "required but ignored," so any string will do.

from openai import OpenAI

client = OpenAI(
    base_url='http://localhost:11434/v1/',
    api_key='ollama',  # required but ignored
)

chat_completion = client.chat.completions.create(
    messages=[{'role': 'user', 'content': 'Say this is a test'}],
    model='qwen3',
)
print(chat_completion.choices[0].message.content)

🔌 OpenAI compatibility is powerful: many libraries and tools support the OpenAI API. Point them at Ollama's /v1 endpoint and you can use local models instead of the cloud. It also works as a fallback when a cloud service goes down. Python and JavaScript also have official libraries (pip install ollama / npm i ollama).

7. Customizing (Modelfile, env vars)

Modelfile: build your own model

A Modelfile is a config file much like a Dockerfile. Start from a base model (FROM, required), add parameters (PARAMETER) and a system prompt (SYSTEM), and you get "your own model" that always behaves the same way. A minimal example:

FROM qwen3
PARAMETER temperature 0.7
PARAMETER num_ctx 8192
SYSTEM """You are a concise assistant. Lead with the key point and keep answers short."""

Save it as a file named Modelfile, then run this in the same folder:

ollama create my-assistant -f Modelfile
ollama run my-assistant

To see the Modelfile behind an existing model, run ollama show --modelfile qwen3. Copying and editing that is the quickest way to start.

Key environment variables

You change how the server behaves with environment variables. ollama serve --help prints the full list. Here are the common ones, with the defaults written in the source code.

Variable What it changes Default
OLLAMA_HOST The address and port the server listens on; 0.0.0.0:11434 lets other devices on your LAN use it 127.0.0.1:11434
OLLAMA_MODELS Where models are stored; use it to move them to a bigger drive macOS ~/.ollama/models, Windows C:\Users\%username%\.ollama\models, Linux /usr/share/ollama/.ollama/models
OLLAMA_KEEP_ALIVE How long a model stays in memory after its last use; -1 keeps it loaded 5m (5 minutes)
OLLAMA_CONTEXT_LENGTH The default context length Set by VRAM (under 24 GiB: 4k / 24–48 GiB: 32k / 48 GiB and up: 256k)
OLLAMA_NUM_PARALLEL How many requests one model handles at the same time 1
OLLAMA_ORIGINS Origins allowed to access it from a browser (comma-separated) 127.0.0.1 and 0.0.0.0
OLLAMA_NO_CLOUD Set to 1 to turn off cloud features (remote inference and web search) and stay local-only Off

How you set them depends on the OS (official FAQ). On Windows, quit Ollama from the taskbar, add a user variable in the environment variables settings, and start Ollama again from the Start menu. On macOS, set it with something like launchctl setenv OLLAMA_HOST "0.0.0.0:11434" and restart the app. On Linux (when running under systemd), do this:

systemctl edit ollama.service
# add one line per variable under [Service]
# Environment="OLLAMA_HOST=0.0.0.0:11434"
systemctl daemon-reload
systemctl restart ollama

⚠️ Watch the context length: the default context length depends on how much VRAM you have, and below 24 GiB it is only 4k. That's too short for long documents or coding agents, and the model silently forgets earlier content without any error. The symptoms and the fix are covered in the Ollama + Cline article.

8. Troubleshooting

Here are the common snags and fixes, up front.

Slow or stalling

If the PROCESSOR column in ollama ps doesn't say "100% GPU", the model doesn't fit in VRAM. Step down a size or use a more heavily quantized build.

Crashes from low memory

Plan on at least 8 GB RAM even for 7B, and about 16 GB for 13B and up. A longer context uses even more, so keep num_ctx to what you need.

API won't connect

Check that the server (the app or ollama serve) is running and nothing else is using port 11434. To connect from another machine, set OLLAMA_HOST to 0.0.0.0.

Model not found

Usually a typo in the model name or size tag. Check the correct name in the official model list.

Finding the logs

On Windows, see server.log in %LOCALAPPDATA%\Ollama; on macOS, ~/.ollama/logs; on Linux, run journalctl -u ollama.

Running out of space on C:

Move model storage to another drive with OLLAMA_MODELS. The install location itself can be changed with OllamaSetup.exe /DIR="d:\some\location".

Summary

Ollama is the fastest way to get into local LLMs. Three takeaways:

  • Set up in minutes: install from the official site, then just ollama run <model>. Very few commands to learn.
  • Choose models by size: stay within your VRAM. When unsure, start at the 7B class and pick a lineage by use case.
  • The API is the real value: the OpenAI-compatible API at localhost:11434 lets you embed it in your own apps and chat UIs—and serve as a cloud fallback.

Start by typing ollama run qwen3. The best way to learn is to run it while checking the differences from the cloud and how to choose a model.

From there, if you want to go on to wiring a model into your editor and having it write code, coding with a local LLM covers configuring Cline and Continue, along with the context length that makes an agent break without ever raising an error.

FAQ

Q. Is Ollama free? Can I use it commercially?

A. Ollama itself is free and open-source. However, each model you run has its own license, and commercial use depends on the model. Check each model's terms before product use (see the licensing section of our model comparison).

Q. Ollama or LM Studio—which is better?

A. For commands, APIs, automation, and embedding into your own apps, Ollama; if you want to start easily with a GUI, LM Studio. Both are free, so when unsure, install both and compare.

Q. Is my data sent externally?

A. As long as you run a model on your own machine, processing stays on your PC; the official FAQ states that Ollama doesn't see your prompts or data when you run locally. But Ollama also offers cloud-hosted models, and if you pick one, your input is processed on Ollama's servers (Ollama says it neither stores it nor trains on it). To stay local-only, set OLLAMA_NO_CLOUD=1.

Q. How do I update or uninstall it?

A. On Windows and macOS, updates download automatically and apply when you click "Restart to update" from the icon. On Linux, re-run the install line. To uninstall on Windows, remove it from the apps list in Settings; on macOS, delete the items listed on the official macOS page. Models stay behind in the .ollama folder, so delete that too if you no longer need them.

Q. Can I use it with existing OpenAI code?

A. Yes. Ollama exposes an OpenAI-compatible API at localhost:11434/v1, so in most cases you only change the endpoint URL and the model name. Handy for switching from cloud to local, or as a fallback.

Q. What kind of PC do I need?

A. As a guide, at least 8 GB RAM for 7B models and 16 GB+ for 13B and up. For comfort, a supported GPU (8 GB+ VRAM) or a Mac with ample unified memory helps. See the hardware requirements article for details.