kmail.at
A colossal neural network lattice suspended in darkness, layers of fine filaments receding away from an empty centre

omar · local ai, measured against your machine

Which AI model runs on your machine?

Omar reads the hardware you are reading this on and tells you which open-weight models will actually run — the memory they need, the speed you can expect, and how to download them. All of it in your browser. None of it uploaded.

104
models scored
23
providers
292
devices covered
09
use cases

the finder

Start with your machine

✓ nothing leaves your device — detection runs in the browser

your machine

detecting…

platform
—
cpu cores
12
system ram
—
gpu
—
vram
—
bandwidth
—

Computed on your device. No request carries these numbers anywhere.

Detecting your hardware…

Three moves, no upload

01

It reads your hardware

Your GPU, CPU cores and memory are read from the browser through WebGL and WebGPU — no install, no extension, no benchmark download. Every figure is labelled as an estimate, because that is what a browser can honestly give.

02

It scores every model

Each of the 104 models is graded against what it just read: the memory it needs at a chosen quantisation, the tokens per second your bandwidth implies, and a fit grade from S to F.

03

It tells you how to run it

Pick a model and Omar hands you the actual commands — Hugging Face, Ollama, LM Studio, llama.cpp, Unsloth — generated for the quantisation you selected, with links to download it or find more.

What a grade is made of

The score is not a vibe. Three measured things, weighted 55 / 35 with a small capability bonus, scaled down when the fit is tight or the model has to be offloaded to system RAM.

Memory
The model at your chosen precision must fit — with headroom. Comfortably inside scores 100; nearly full scores near zero.
Speed
Memory bandwidth divided by the bytes read per token, times an efficiency factor (0.70 discrete GPU, 0.65 Apple Silicon, 0.40 mobile).
Capability
A small bonus for larger models, so a big model that fits beats a small one that fits easily.

Grades, at a glance

  • SRuns great — comfortable headroom, fast
  • ARuns well
  • BAcceptable — less margin
  • COn the edge — expect offloading
  • DBarely runs
  • FHeavier than this machine

quality against size · a 7B model

  • F16~13 GB100%
  • Q8_0~6.7 GB~99%
  • Q6_K~5.3 GB~95%
  • Q4_K_M~3.9 GB~88%
  • Q2_K~2.5 GB~60%

It ends in a command you can paste

Every model page carries a “how to run it” section: the download and run commands for six runtimes, generated for the precision you chose, each checked against the tool’s own CLI.

What is inside

Models
104 open-weight models, 23 providers, every published quantisation with memory and disk size per level.
Discrete GPUs
214 cards — NVIDIA, AMD and Intel — with VRAM and memory bandwidth, the two figures that decide local inference.
Apple Silicon
20 M-series tiers with unified memory and bandwidth.
Mobile & tablets
40 Adreno, Mali and Tensor GPUs, with the platform-specific memory rules applied.
Devices
292 named devices, each showing how many catalogue models run well on it.
Runtimes
6 ways to actually run the model, each with verified commands and links.