kmail.at
← learning

technologies · difficulty ◆◆

Ollama — run language models from your terminal

Docker grammar for language models: pull, run, ps, stop — local AI in five commands.

You already know how to run a language model on your own hardware. You learned it in 2013, from Docker: pull an image, run it, list what is running, stop it when you are done. Ollama borrowed that grammar wholesale, pointed it at model weights instead of containers, and the result is the friendliest command line in AI.

2026-10-04 · 12 min read

$ ollama

What it is

Ollama is a free, open-source command-line runner for open-weight language models, built in Go around the llama.cpp inference engine. It packages model download, hardware detection, and chat into a single binary: type ollama run gemma4 and it fetches the weights, picks CUDA, Metal, Vulkan, or plain CPU without asking you, and starts a conversation in your terminal. The whole design is Docker for language models — pull a model like an image, run it like a container, list running models with ps, stop them with stop.

Why it matters

Before Ollama, running an open model meant compiling llama.cpp, converting weights to GGUF by hand, learning quantization levels, and debugging CUDA versions. The founders said it plainly in a 2026 interview: open models in 2023 were geared toward researchers, not programmers. Ollama collapsed all of that into one command. It also starts a local REST API on port 11434 automatically, so anything built against the OpenAI SDK can point at localhost:11434 and change nothing but the base URL. And because everything runs on your machine, your prompts stay on your machine — the strongest privacy mode there is.

The model zoo

Models live in a library on ollama.com, and tags work like image tags: a bare name is the default variant, a colon picks a size or quantization. gemma4:e2b is the ~7 GB instruction model the quickstart recommends; smollm2:360m is a 725 MB minnow that runs on nearly anything with a pulse. Third parties publish their own, and the same colon syntax picks a cloud-hosted model — glm-5.3-flash:cloud runs on the company's hardware while behaving exactly like a local one in every command. The command surface itself is small: list, run, pull, push, ps, stop, cp, rm, show, create, and a couple of account commands.

Example

$ ollama list
NAME                       ID              SIZE      MODIFIED     
smollm2:360m               297281b699fc    725 MB    3 days ago

Your local model shelf: NAME, ID, SIZE on disk, MODIFIED. The 360m tag means 360 million parameters — small enough to run on a decade-old laptop. The ID is a git-style content hash; two entries sharing an ID are the same weights. A cloud-pointer model shows a dash in SIZE, because it lives in the registry, not on your disk.

$ ollama pull smollm2:360m
pulling manifest 
pulling 7d23be4097d6: 100% ▕██████████████████▏ 725 MB
pulling fbacade46b4d: 100% ▕██████████████████▏   68 B
pulling d502d55c1d60: 100% ▕██████████████████▏   675 B
verifying sha256 digest 
writing manifest 
success 

Layers, exactly like docker pull: the weights blob, the template, the license, each with its own digest, verified by sha256. Storage is content-addressed, so a variant that shares most layers with a model you already have downloads in seconds.

$ ollama run smollm2:360m "Why is the sky blue? Answer in one short sentence."
The sky appears blue because of a phenomenon called Rayleigh scattering, where shorter wavelengths of light are scattered more than longer wavelengths, giving the impression of blue.

One-shot mode: prompt as argument, answer to stdout, exit. On a fresh model the first run downloads it first — the official quickstart demos exactly this with gemma4:e2b — so pull and run are the same step when convenient. No API key, ever.

$ ollama run smollm2:360m "Name one thing octopuses can do." --verbose
One thing an octopus can do is change its skin texture, color, and shape to adapt to their surroundings, prey, or predators.

total duration:       1.267192095s
load duration:        1.376343ms
prompt eval count:    41 token(s)
prompt eval duration: 70.878ms
prompt eval rate:     578.46 tokens/s
eval count:           28 token(s)
eval duration:        1.190171s
eval rate:            23.53 tokens/s

The honest stopwatch. eval rate is generation speed in tokens per second, prompt eval rate is how fast the model reads your input, load duration tells you whether it was already sitting in memory. Benchmark models or hardware with these numbers, never with vibes.

$ ollama run smollm2:360m "Give me a JSON object with keys 'animal' and 'sound' for a cat." --format json
{
  "animal": "cat",
  "sound": "meow"
}

--format json steers the output into a JSON object — this 360-million-parameter model nailed the schema on the first try. Pair it with --think=false on reasoning models when you want the answer without the visible thinking.

$ ollama run smollm2:360m
>>> Send a message (/? for help)

No prompt argument drops you into the interactive REPL: /bye quits, /? lists the slash commands, /set system changes the system prompt mid-session, /show info prints the model card. Conversation memory lives for the session only — exit and it is gone.

$ ollama ps
NAME             ID              SIZE      PROCESSOR    CONTEXT    UNTIL              
smollm2:360m     297281b699fc    919 MB    100% CPU     4096       4 minutes from now

The running-model table. PROCESSOR shows where the model landed (100% CPU on this box, GPU+CPU splits on mixed systems), CONTEXT is the loaded context length, UNTIL counts down the keepalive — 5 minutes by default — after which the memory is freed. This is also your first stop when disk or VRAM looks surprisingly busy.

$ ollama run no-such-model-xyz "hi"
pulling manifest ⠋ pulling manifest ⠙ ... 
Error: pull model manifest: file does not exist

The classic first error. A typo, or a name that does not exist in the library, fails after a few spinner turns of manifest checking. Note what Ollama tried first: it attempted to pull it before complaining. When this happens, check spelling against ollama.com/library or run ollama list to see what you actually have.

$ printf 'FROM smollm2:360m\nSYSTEM """You are a pirate. Reply in pirate speak."""' > Modelfile
ollama create pirate-helper -f Modelfile
gathering model components 
using existing layer sha256:7d23be40...
creating new layer sha256:f1b3345a34bb...
writing manifest 
success 

A Modelfile is a Dockerfile for a personality. FROM picks the base weights, SYSTEM sets the standing instruction, PARAMETER lines tune generation settings — and ollama create bakes it into a new entry. Real output, abridged: three layers reused from the base, two new ones for the recipe. ollama show pirate-helper then prints back your system prompt, temperature, and the base model's stats.

$ ollama cp pirate-helper pirate-helper-backup && ollama rm pirate-helper-backup
copied 'pirate-helper' to 'pirate-helper-backup' 
deleted 'pirate-helper-backup' 

cp duplicates a model entry — it copies the manifest, not 725 MB, so both names share the same weights — and rm removes entries. rm on a model you created releases its unique recipe layers; the base model stays on disk because the original still references it. That is why rm so often frees less space than people expect: layers go only when nothing points at them anymore.

Common flags

ollama serve
start the server process itself (auto-starts from the desktop app; needed on a headless Linux box)
ollama run M [PROMPT]
chat interactively, or one-shot with a prompt argument; auto-pulls missing models
ollama run M --verbose
print per-response timing: load, prompt eval, eval counts and rates
ollama run M --format json
force JSON output; also accepts a JSON schema
ollama run M --think=true|false|high
toggle reasoning depth on thinking models; --hidethinking hides the trace
ollama run M --keepalive 30m
keep the model in memory for a duration instead of the default 5m
ollama pull / push M
download from / upload to the registry layers-first with sha256 verification
ollama list
every local model: name, ID hash, disk size, modified date
ollama ps
currently loaded models: processor placement, context, keepalive countdown
ollama stop M
unload a model immediately instead of waiting out the keepalive
ollama show M [--modelfile|--parameters|--license]
X-ray a model: architecture, params, quantization, template, system prompt
Modelfile + ollama create M -f
build a custom model: FROM a base, add SYSTEM and PARAMETER lines, quantize with -q
ollama cp A B
copy a model entry — cheap, useful for A/B versions of a recipe
ollama rm M...
delete local models, one or several at once
ollama create M --quantize q4_K_M
build a compressed 4-bit variant at create time
OLLAMA_HOST
env var to reach a remote Ollama server, or move the serve port off 11434
ollama launch claude|codex|hermes
wire terminal coding agents to Ollama models in one step
ollama signin / signout
account session for the registry and cloud models

History

Kitematic, sold to Docker, round two

Jeffrey Morgan and Michael Chiang built Kitematic, a GUI that made Docker approachable on macOS. Docker acquired it in 2015 and folded it into Docker Desktop, a tool used by millions. In 2023 the founders noticed the newly released open models were, in Morgan's words, geared toward researchers rather than programmers, and rebuilt the Kitematic trick at the command line: wrap hard infrastructure in a familiar interface. The Docker-inspired grammar of ollama pull and ollama run was the entire product insight.

The borrowed engine

Ollama's actual inference is llama.cpp, Georgi Gerganov's C/C++ runtime from March 2023 that made LLaMA weights run on consumer hardware. The repo was created on June 26, 2023, and Ollama launched publicly in July 2023 on top of it, with hardware autodetection and a friendly registry layered over the engine. The full stack story continues today: GGUF single-file model format, quantization from Q2_K to Q8_0, and the same engine that also carries 120k+ stars of its own on GitHub.

2026: cloud pointers and mainstream scale

Ollama has since grown past 182,000 GitHub stars and closed a $65M Series B in July 2026 ($88M total funding). Per the announcement, roughly 8.9M developers use it, and the community has built more than 67,000 integrations on top of the platform — Open WebUI alone, the self-hosted chat frontend, talks to Ollama through the same localhost API you just learned. The 0.3x releases added cloud models, which appear in ollama list with a dash for size: same commands, weights in the registry.

Fun facts

Pros & cons

pros

  • + One command from zero to running model, with no API keys for local use
  • + Layered, content-addressed storage: variants share blobs and pull in seconds
  • + ps / stop / keepalive give container-grade control over memory
  • + OpenAI-compatible API on localhost:11434 — most OpenAI SDK code ports by changing one URL
  • + Modelfile + create makes custom personalities a two-line file
  • + MIT-licensed core, llama.cpp underneath, honest --verbose timing on every response

cons

  • − Model quality on small hardware is bounded by physics: a 725 MB model writes like a 725 MB model
  • − No built-in conversation memory across sessions — the REPL forgets everything on /bye
  • − Cloud models blur the "local" story: the :cloud tag phones home more than the branding suggests
  • − The REST API is a separate surface (curl, not CLI flags) — this tutorial stays on one side of that wall

Takeaways

  1. 1Install once, then ollama run smollm2:360m is your hello-world; upgrade to gemma4:e2b when 725 MB stops being funny and starts being limiting.
  2. 2ollama list tells you what is on disk, ollama ps tells you what is in memory, ollama stop unloads on demand — that trio is 90% of daily hygiene.
  3. 3--verbose is your benchmark harness: eval rate tokens/s across models, same prompt, on your hardware, no marketing numbers.
  4. 4Build a Modelfile with FROM + SYSTEM + PARAMETER, run ollama create, and your reusable persona is one ollama run away.
  5. 5Point an OpenAI-SDK app at http://localhost:11434 when the CLI is not enough — same models, same server, structured calls.

Related commands

← all learning