kmail.at
← tools

osama

v0.1.0 · beta · ml

the full llama.cpp toolbox as a desktop GUI, plus MLX on Apple silicon — no terminal required

Osama's dashboard: engine state, library and tool coverage
The dashboard — engine state, the library, and tool coverage, with the agent harness counted deliberately apart from the llama.cpp binaries.
The model library and Hugging Face discovery on one page
One page for the models you have and the models you could fetch: the one you just downloaded is right there in the library.
Serving a model as an OpenAI-compatible API
Serve a model and point anything at it. The endpoint, the health check and the curl line are all on the page.
The quantize tool form, built from the real CLI flags
Every tool form is built from the real CLI flags — the help text, the defaults, and the exact command it will run.
Installed llama.cpp builds and recent releases
Install and switch official builds; each one is enumerated down to the tools it actually ships.

about

Osama is a desktop studio for running models on your own machine. It does

not reimplement llama.cpp — it installs the official ggml-org release

binaries, runs them, and puts a form in front of every tool they ship.

The current build carries twenty-four binaries, sixteen of them driven

through generated front-ends: quantize with an importance matrix, split

and merge GGUF files, merge LoRAs, benchmark throughput, evaluate, and

inspect. Each field is declared from the flags the binary actually

accepts, so the form explains itself, the command it builds is the

command you would have typed, and the argv is passed as an array rather

than a shell string.

Models come from wherever they live — Hugging Face, ModelScope, CivitAI,

Ollama's registry, or a .gguf you pasted — and downloads resume rather

than restarting. Once a model is on disk you can chat with it, or serve

it as an OpenAI- and Anthropic-compatible API and point any client at

it. A local agent harness sits alongside: thirty-seven tools the model

may call, kept deliberately apart from what the llama.cpp build ships,

jailed to a workspace and gated by approval on anything that writes.

On Apple silicon a second engine runs beside the first: MLX, driving

Apple's mlx-lm, sharing the models page, the chat and the API. macOS

(Apple silicon or Intel), Linux and Windows. MIT.

what it does

The real binaries
Installs and switches official ggml-org llama.cpp builds rather than shipping a fork, and enumerates every tool a build actually contains.
Models from anywhere
Hugging Face, ModelScope, CivitAI, Ollama's registry, or a pasted .gguf. Resumable, multi-file, and honest about what a file is.
Sixteen tool front-ends
Quantize, imatrix, split and merge, LoRA merge, bench, evaluate, inspect — every form generated from the flags the binary accepts.
Chat that keeps its place
Fork, edit, regenerate and continue a turn; attachments read as text, PDFs extracted. A turn survives leaving the page.
A server in one click
OpenAI- and Anthropic-compatible, with API keys, parallel slots, Jinja templates and Prometheus metrics — and the curl line to paste.
A local agent harness
Thirty-seven built-in tools across filesystem, network and state, jail-rooted to a workspace and gated by approval on anything mutating.
MLX on Apple silicon
Apple's mlx-lm as a second engine — same models page, same chat, same API — and gated off entirely where it cannot run.
Desktop, not a browser tab
A Tauri 2 shell on macOS, Linux and Windows. Your models, your machine, nothing uploaded.

synopsis

git clone https://github.com/mokmail/osama && cd osama && npm install && npm run build && npm start

changelog

  • v0.1.0initial release: engine installer, HF discovery + resumable downloads, chat, llama-server manager, quantize/imatrix/split, llama-bench, headless CLI, Tauri 2 shell