
omar · local ai, measured against your machine
Which AI model runs on your machine?
Omar reads the hardware you are reading this on and tells you which open-weight models will actually run — the memory they need, the speed you can expect, and how to download them. All of it in your browser. None of it uploaded.
- 104
- models scored
- 23
- providers
- 292
- devices covered
- 09
- use cases
the finder
Start with your machine
✓ nothing leaves your device — detection runs in the browser
your machine
detecting…
- platform
- —
- cpu cores
- 12
- system ram
- —
- gpu
- —
- vram
- —
- bandwidth
- —
Computed on your device. No request carries these numbers anywhere.
Detecting your hardware…
Three moves, no upload
01
It reads your hardware
Your GPU, CPU cores and memory are read from the browser through WebGL and WebGPU — no install, no extension, no benchmark download. Every figure is labelled as an estimate, because that is what a browser can honestly give.
02
It scores every model
Each of the 104 models is graded against what it just read: the memory it needs at a chosen quantisation, the tokens per second your bandwidth implies, and a fit grade from S to F.
03
It tells you how to run it
Pick a model and Omar hands you the actual commands — Hugging Face, Ollama, LM Studio, llama.cpp, Unsloth — generated for the quantisation you selected, with links to download it or find more.
What a grade is made of
The score is not a vibe. Three measured things, weighted 55 / 35 with a small capability bonus, scaled down when the fit is tight or the model has to be offloaded to system RAM.
- Memory
- The model at your chosen precision must fit — with headroom. Comfortably inside scores 100; nearly full scores near zero.
- Speed
- Memory bandwidth divided by the bytes read per token, times an efficiency factor (0.70 discrete GPU, 0.65 Apple Silicon, 0.40 mobile).
- Capability
- A small bonus for larger models, so a big model that fits beats a small one that fits easily.
Grades, at a glance
- SRuns great — comfortable headroom, fast
- ARuns well
- BAcceptable — less margin
- COn the edge — expect offloading
- DBarely runs
- FHeavier than this machine
quality against size · a 7B model
- F16~13 GB100%
- Q8_0~6.7 GB~99%
- Q6_K~5.3 GB~95%
- Q4_K_M~3.9 GB~88%
- Q2_K~2.5 GB~60%
It ends in a command you can paste
Every model page carries a “how to run it” section: the download and run commands for six runtimes, generated for the precision you chose, each checked against the tool’s own CLI.
- Hugging FaceThe source of the weights. Download the exact files, or use the same CLI to fetch a different quantisation of the same model.
- OllamaThe least friction: one command runs the model, and Ollama manages the quantisation and the server for you.
- LM StudioA desktop app with a built-in model browser and an OpenAI-compatible local server — the option with no terminal at all.
llama.cppThe engine most of the others wrap. Most control over context, offloading and sampling — and the one to reach for when a quantisation is not in anybody’s registry.
UnslothFor going past inference: fine-tune or convert this model on a single GPU, and export your own quantisations.
Pi · kmailaiRun a local model as a coding agent rather than a chat box. This is the Pi engine, packaged as kmailai, and it talks to anything with an OpenAI-compatible endpoint.
What is inside
- Models
- 104 open-weight models, 23 providers, every published quantisation with memory and disk size per level.
- Discrete GPUs
- 214 cards — NVIDIA, AMD and Intel — with VRAM and memory bandwidth, the two figures that decide local inference.
- Apple Silicon
- 20 M-series tiers with unified memory and bandwidth.
- Mobile & tablets
- 40 Adreno, Mali and Tensor GPUs, with the platform-specific memory rules applied.
- Devices
- 292 named devices, each showing how many catalogue models run well on it.
- Runtimes
- 6 ways to actually run the model, each with verified commands and links.