osama
v0.1.0 · beta · mlthe full llama.cpp toolbox as a desktop GUI, plus MLX on Apple silicon — no terminal required





about
Osama is a desktop studio for running models on your own machine. It does
not reimplement llama.cpp — it installs the official ggml-org release
binaries, runs them, and puts a form in front of every tool they ship.
The current build carries twenty-four binaries, sixteen of them driven
through generated front-ends: quantize with an importance matrix, split
and merge GGUF files, merge LoRAs, benchmark throughput, evaluate, and
inspect. Each field is declared from the flags the binary actually
accepts, so the form explains itself, the command it builds is the
command you would have typed, and the argv is passed as an array rather
than a shell string.
Models come from wherever they live — Hugging Face, ModelScope, CivitAI,
Ollama's registry, or a .gguf you pasted — and downloads resume rather
than restarting. Once a model is on disk you can chat with it, or serve
it as an OpenAI- and Anthropic-compatible API and point any client at
it. A local agent harness sits alongside: thirty-seven tools the model
may call, kept deliberately apart from what the llama.cpp build ships,
jailed to a workspace and gated by approval on anything that writes.
On Apple silicon a second engine runs beside the first: MLX, driving
Apple's mlx-lm, sharing the models page, the chat and the API. macOS
(Apple silicon or Intel), Linux and Windows. MIT.
what it does
- The real binaries
- Installs and switches official ggml-org llama.cpp builds rather than shipping a fork, and enumerates every tool a build actually contains.
- Models from anywhere
- Hugging Face, ModelScope, CivitAI, Ollama's registry, or a pasted .gguf. Resumable, multi-file, and honest about what a file is.
- Sixteen tool front-ends
- Quantize, imatrix, split and merge, LoRA merge, bench, evaluate, inspect — every form generated from the flags the binary accepts.
- Chat that keeps its place
- Fork, edit, regenerate and continue a turn; attachments read as text, PDFs extracted. A turn survives leaving the page.
- A server in one click
- OpenAI- and Anthropic-compatible, with API keys, parallel slots, Jinja templates and Prometheus metrics — and the curl line to paste.
- A local agent harness
- Thirty-seven built-in tools across filesystem, network and state, jail-rooted to a workspace and gated by approval on anything mutating.
- MLX on Apple silicon
- Apple's mlx-lm as a second engine — same models page, same chat, same API — and gated off entirely where it cannot run.
- Desktop, not a browser tab
- A Tauri 2 shell on macOS, Linux and Windows. Your models, your machine, nothing uploaded.
synopsis
git clone https://github.com/mokmail/osama && cd osama && npm install && npm run build && npm startchangelog
- v0.1.0initial release: engine installer, HF discovery + resumable downloads, chat, llama-server manager, quantize/imatrix/split, llama-bench, headless CLI, Tauri 2 shell