ai
System One Goes Local: Five Models, Eighteen Days, One Wire Format
Three weeks after Jev launched as the only System One model, the category has five new open models on Ollama (Nimble 9B, Tev1 4B/0.8B, Clef 27B, Clef-Flash 9B), a native /v1/systemone endpoint in Ollama 0.35, llama.cpp support, and a three-way benchmark war. The full map of the wave, with the honest reading of every number.

System One Goes Local: Five Models, Eighteen Days, One Wire Format
Three weeks ago there was exactly one System One model — TypeSafe’s Jev, a hosted API that could not write a sentence. Today the category has a benchmark war, a package manager-shaped runner, a home in llama.cpp, and a whole shelf of models you can pull down and run on your own hardware with one command. This is the map of that wave: what a decision model actually is, who has shipped one, how they compare, and what the new open ecosystem can and cannot do.
The three-week-old category, explained
On September 15, 2026, TypeSafe AI came out of two years of stealth with $40 million in seed funding and an idea that sounds almost rude: the most useful model for most software is one that never generates text.
A large language model answers a question by writing prose one token at a time, leaving your application to parse meaning back out of a string. A System One model does the opposite. You send a block of program state plus a set of typed questions — and it returns typed answers with probabilities, evaluated in a single parallel pass. No parsing step, no schema repair, no prose to defend in a code review. The category name is a borrow from Daniel Kahneman: System 1 is fast, intuitive judgment; System 2 is slow deliberation, and that is roughly what your chat model is for.
TypeSafe trained Jev with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD), and the property that matters is mathematical: under a strictly proper scoring rule, the only optimal strategy is reporting honest probabilities. The model is not asked nicely to be calibrated — it is scored in a way that makes honesty the winning move. Jev exposes three typed primitives: choice (pick one option from a list you supply), score (place the state on an ordered scale you define), and noul (the probability that a proposition is true). Pricing is $0.042 per million input tokens with output free — TypeSafe’s own words are "too cheap to meter" — and latency of 70ms to 500ms. The demand came fast and visibly: the waitlist cleared within days of launch, and by late September Jev was callable through Cloudflare Workers AI, Vercel’s AI Gateway (which reported it reaching nearly 13 percent of paid teams within 24 hours, twice the share of the GPT-5.6 family), Netlify, LangChain’s TypeSafeClassifier and AutoMode middleware, Pydantic AI, and five independent Elixir clients.
"Think of Jev as a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out."
Real-world reports landed in the same window. A Vercel engineer measured a safety classifier running five to 18 times faster than the LLM it replaced; OpenChamber’s analysis of 12,759 launch posts put user-reported speedups at a median of 7x, cost savings at a median of 30x, and latency at a median of 76ms — meaningfully below the 193.6x headline but still a different genus of speed than a chat completion. And Armin Ronacher, whose opinion on API design carries some weight in this industry, noted the design "delegates the hallucination problem a little bit to the user," who now decides whether a 50 percent probability is worth acting on.
Why this wave is different from every other model race
One sentence: the decision layer is becoming infrastructure, and infrastructure wants to be open.
Jev is excellent — the docs are unusually honest about its limits, the pricing is nearly free, and the ecosystem adopted its wire protocol as the standard. But it is a hosted endpoint behind an API key: no weights, no on-premise option, no published parameters. For three weeks that was the whole story. What happened next is the part that should excite anyone who runs their own hardware: in under three weeks, the category produced open-weights replicas in every size class, native support in two of the most-deployed local inference tools on earth (Ollama and llama.cpp), a Cloudflare front-end with vision-capable 27B weights, and multiple independent leaderboards keeping score. The wire format converged so fast that switching between the closed frontier API and a model on your own GPU is, in most cases, a one-line base-URL change.
If you have been waiting for the "open weights catch-up moment" before taking this category seriously — this is it. In building my own decision-layer integration over the jevai.org API — a routing guard where no verdict means no execution — the same question template ran against both surfaces unchanged, and the open ecosystem inherited exactly that portability: the official typesafe-sdk client works against Ollama, Laya, lev, decider, Kev, and Liquid’s d1 with only a base-URL change, because every one of them adopted TypeSafe’s /v1/systemone wire format. Your thresholds, your question schemas, your escalation logic — all portable.
Ollama adds decision models natively — the headline
On September 29, 2026, Ollama shipped 0.35, one of those releases that quietly redefines what the tool is for: native decision-model support through a new /v1/systemone endpoint, wire-compatible with Jev. ollama pull, curl http://localhost:11434/v1/systemone, done — the same request shape and response schema as the hosted API, zero per-decision cost, and no network round-trip.
ollama pull nimblecurl http://localhost:11434/v1/systemone -d '{
"model": "nimble",
"state": {"ticket": "I was charged twice. Please refund the extra payment."},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team should handle this ticket?",
"criteria": {
"billing": "Payments and refunds",
"technical": "Bugs and integrations",
"other": "None of the above"
}
},
"refund": {
"type": "noul",
"instructions": "Does the customer explicitly ask for a refund?"
},
"urgency": {
"type": "score",
"instructions": "How urgent is this ticket?",
"criteria": ["Routine", "Soon", "Urgent"]
}
}
}'The response carries per-option probabilities, a confidence value, and literal token accounting:
{
"model": "nimble",
"answers": {
"team": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.985, "technical": 0.012, "other": 0.003},
"confidence": 0.922},
"refund": {"type": "noul", "noul": 0.997},
"urgency": {"type": "score", "score": 0.815,
"legend": {"0": "Routine", "1": "Soon", "2": "Urgent"},
"probabilities": {"0": 0.378, "1": 0.429, "2": 0.193},
"confidence": 0.046}
},
"usage": {"input_tokens": 841, "output_tokens": 4}
}Four output tokens for three typed decisions. That usage line is the whole thesis of the category in miniature.
The launch lineup is three models, and it has since grown to five; more, including cloud-served variants, are explicitly teased. The fine print matters, and Ollama’s docs are refreshingly blunt about it: requests without images must fit in 64 KiB; the local schema caps choice questions at 2 to 26 named options; and confidence "measures how strongly the model favors one answer over the others... A higher value does not guarantee the answer is correct."
The new Ollama lineup: five models, three families
Nimble — Bespoke Labs’ 9B runner-up that beats the winner on some days
Nimble is a 9B decision model from Bespoke Labs, fine-tuned from Qwen3.5-9B, under Apache 2.0. Its readme is a tutorial in how this class of model works: answers are scored as single letter codes ("one token per answer" — there is no generated JSON to parse), it reads the prompt once per question and scores the answer tokens directly ("there’s no reasoning step. This is what makes it fast"), it was trained on contrastive pairs whose members differ by one fact that flips the answer, and it accepts up to 64 questions about the same text in one call. In Ollama’s own launch demo on an M5 Max, Nimble 9B averaged 91ms per decision.
The numbers deserve the honest reading. Ollama’s blog chart — mean accuracy across 13 public datasets with human labels, 3,880 decisions — puts Nimble at 75.7 percent against Jev 1.13’s 76.0, with the Jev figure coming from Bespoke Labs’ own published run of the same decisions. On Bespoke’s held-out set, Jev is clearly ahead: 93.21 to 90.12 percent agreement. And on Ollama’s launch chart, Nimble loses the reasoning-panel 85.7 to 76.0 while leading Jev 95.6 to 89.0 on factual recall (TREC). Translation: a local 9B model you can pull and run is trading a few points of average accuracy against the hosted frontier for 91ms local latency, zero marginal cost, and weights you own — and it wins entire categories some days.
Tev1 — Together AI’s experimental pair, trained for $17 and open in every way but the license
Tev1 is a family of experimental decision models from Together AI, fine-tuned from Qwen3.5 in two sizes: 4B (the accurate one) and 0.8B (an 812MB download for tight memory budgets). Its provenance statement is unusually thorough — 37,840 training examples across MultiNLI, BoolQ, Banking77, AG News, SST-5, programmatic policies, routing, and a research taxonomy — plus its own sharp sentence about genealogy: "Tev1 takes its inspiration from Jev. None of its training data came from Jev." Together published a full companion piece on fine-tuning your own Jev-like model for $17, which is the number that turns this from a vendor story into a hobbyist story.
The caveats are documented rather than discovered: Tev1 was trained on 2 to 24 options and stays in that range; it carries a practical context of about 2,000 tokens (the longest training example is ~1,500); in a regular chat it tends to reply in prose, because the decision behavior only appears through /v1/systemone; and "Together AI hasn’t fully tested prompt injection, languages other than English, calibration, or how it handles inputs unlike its training data." One licensing nuance belongs in any procurement note: the dataset builders and training scripts are MIT licensed, but the Ollama page does not state a license for the weights themselves — check the GitHub repo before shipping anything commercial. On Ollama’s 13-dataset chart Tev1 4B scores 73.3 and the 0.8B scores 63.5.
Clef and Clef-Flash — Cloudflare’s 27B eyes, and the fastest decision model anyone has measured
Days after the Ollama support landed, Cloudflare shipped the category’s biggest surprise: Clef, a 27B decision model fine-tuned from Qwen3.8-27B, and Clef-Flash, a 9B sibling on Qwen3.5-9B — both Apache 2.0, both "fully compatible with the Jev and System One APIs," and both multimodal. Add base64-encoded images (screenshots, receipts, forms, photos) to the request and they are scored jointly with the text. Clef needs Ollama 0.35.1 or later and 18GB of disk at its latest tag; Flash takes 11GB.
The performance claims are the boldest in the category. Clef’s card says it "tops the Decision Index, Cloudflare’s leaderboard of decision models, and beats Jev on most of its suite"; its comparison table has Clef at 94.2 on BANKING77 against Jev’s 79.7, with median latency 209.3ms against Jev’s 524.1. And Clef-Flash posts a median latency of 38.8ms — the lowest of any decision model in Cloudflare’s Decision Index runs, roughly thirteen times faster than the hosted frontier model while answering questions about images. The usual caveat applies — this is Cloudflare’s own leaderboard and their own run — but the numbers are specific, public, and checkable, which already puts the category’s marketing standard above most.
The wider open ecosystem — a field guide
The replicas and the researchers
Kev is Jared Palmer’s Apache-2.0 family (0.8B/4B/9B/27B on Qwen3.5/3.8 bases), built on the architecture the community reverse-engineered for Jev: LoRA plus a pointer head scoring option letters at their own answer slot. The Kev-0.5B prototype, trained on a laptop on September 17 — two days after Jev’s launch — proves the mechanism with 9.3M trainable parameters (1.9 percent of the backbone). Kev 1.0 is the production family; Kev’s own evals put Kev-27B within roughly one point of Jev on out-of-domain accuracy (0.851 vs 0.857), with the honest caveat nobody can currently test: nobody knows what Jev was trained on, so this is not a controlled comparison.
lev, from Interfaze, is an Apache-2.0 LoRA adapter on Qwen3.5-4B (about 200MB on top of an 8GB base) you serve with lev serve. Its self-benchmark is the honesty gold standard of the month: on the 3,880-item S1Bench harness Jev scores macro 0.761 to lev’s 0.689, with lev publishing exactly where it is badly beaten (minimal-edit contradiction pairs, consistency judgment). It is also the clearest latency reality-check in the ecosystem: lev’s engine computes a short request in 69ms, but from a laptop the hosted Jev answered in a 335-346ms median while lev on a cloud H100 took 414-654ms — speed is not automatically the open model’s advantage end-to-end.
decider (Mapika) is the weight-class king of open families — 0.8B, 2B, 4B, 12B, and a 35B mixture-of-experts on Qwen3.5 bases, all Apache 2.0, iterating so fast the 2B went v8 to v11 inside a fortnight. Its card is the category’s promise in one line: "no decoding, no parsing and no output outside the options you defined." Winnow-12B (EldanRing) is the contrarian — a Gemma 4 12B fine-tune shipped as GGUF only, served by a patched llama.cpp server that speaks /v1/systemone and chat from the same loaded model, with vision and 64K context on a 16GB card. On the public JevBench subset, Winnow-12B Q8 matched hosted Jev at 85.7 percent.
The runners, the routers, and the hosted challengers
Ollaya is what you reach for when the decision models are the whole point: an Apache-2.0 community runner with Ollama-style pull/run/serve, its own daemon on port 11435 (next door to Ollama’s 11434), sixteen model families including Laya, Kev, decider, Winnow, and the Clef pair added on October 2 — and it never re-hosts weights; every model is pinned to a commit and verified by sha256, pulled from its author’s repository. Liquid AI’s d1 is the hosted challenger: announced September 29 under d1:free, zero output tokens, from the Liquid Fourier-network lineage, and self-reported as "the first model to outperform Jev on Hugging Face’s Decision Index" — on Liquid’s own run of the board, a claim the rest of the field is not obliged to concede.
And the ecosystem’s scoring layer is already dense. Bespoke’s S1Bench pins the same 3,880 decisions behind every model card. Hugging Face’s Jev Decision Index scored 70 open reproductions across 43 benchmarks at 120,000 decisions per model in its 0.2.1 revision. JevBench — one person’s hobby project, its servers paid out of pocket, MIT-harness — scores 95 systems on a geometric mean of intelligence, calibration, speed, and cost, and its board reads like a thriller: as of v1.4.2.2, an open 4B model (Imajev-4B, 67.37) leads, Jev 1.13.0 sits fourth (63.29), and open takes three of the top five. On that weighted composite, Laya — the open hero of our September deep-dive on Laya vs Jev — ranks 43rd, a reminder that raw speed and honest calibration drag composites when the intelligence axis collapses on hard tiers.
How to choose — the practical decision table
The genuinely useful takeaway is not "who wins" but "which deployment shape fits which constraint." All of these speak the same request format:
- Lowest friction to try today: Ollama +
ollama pull nimble. Apache 2.0, decent accuracy, no network, no API key. - Tightest memory budget: Tev1 0.8B (812MB download), accepting a real accuracy drop.
- Highest local accuracy: Clef 27B — if you can host 18GB and want vision-capable decisions (Ollama 0.35.1+).
- Latency-critical local decisions: Clef-Flash (38.8ms median in Cloudflare’s runs) or Nimble on Apple Silicon (91ms/decision on M5 Max).
- CPU-only / always-on box: Laya’s 421M encoder — the open pick from our September Laya-vs-Jev deep-dive — at a few milliseconds per question via llama.cpp.
- One laptop train-your-own: Kev (per the Jev-Architecture-Unmasked recipe) or Together’s $17 Tev1 recipe.
- Maximum accuracy, no infrastructure: Jev or d1’s free tier — and the hosted option now has real, documented failure modes rather than none.
The benchmark war itself is the caution. Every vendor ranks itself first somewhere — Jev on TypeSafe’s workflow evals and Bespoke’s held-out set; Clef on Cloudflare’s Decision Index; d1 on Liquid’s run of the Decision Index; open models on JevBench’s composite. Third-party corroboration is days old. When a four-billion-parameter model "beats" a frontier model, the correct response is neither applause nor dismissal — it is checking which benchmark, who ran it, and whether the same harness scored the other side.
Running the new models yourself: the honest checklist
- The local stack is young. Decision model support landed in Ollama 0.35 about a week ago and in llama.cpp on October 2; expect churn. Ollama’s library pages state decision models are not yet in the CLI’s interactive flows or some language libraries — API and SDK access first.
- The API quirks are documented, not hidden. Ollama: 64 KiB request cap without images, 2-26 options per
choice, confidence "is not the chance that the answer is right," input never truncated. Nimble: probabilities on your data may not match the readme’s — test any threshold locally first. - Calibration is a measurement, not a guarantee. TypeSafe’s own docs: "Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct." Plan the escalation path for every sub-threshold decision before it runs unattended.
- Read the jaggedness page for whichever hosted model you use. Jev’s documents failures: unreliable counting and arithmetic, date comparison as literal text, accuracy decay on large noisy state, and adversarial content that "can move the answer." Keep arithmetic in code, filter state, pin a versioned model id, and treat a wrong-but-valid typed verdict as the default expected failure mode.
Go deeper
- Ollama now supports Jev-style decision models —
https://ollama.com/blog/ollama-now-supports-jev-style-decision-models - Ollama System One API reference —
https://docs.ollama.com/api/systemone - Ollama decision capability guide —
https://docs.ollama.com/capabilities/decision - Nimble (Ollama library) —
https://ollama.com/library/nimble - Tev1 (Ollama library) —
https://ollama.com/library/tev1 - Clef (Ollama library) —
https://ollama.com/library/clef - Clef-Flash (Ollama library) —
https://ollama.com/library/clef-flash - How to train your own Jev for $17 (Together AI) —
https://www.together.ai/blog/how-to-train-your-own-jev - Clef decision models announcement (Cloudflare) —
https://blog.cloudflare.com/clef-decision-models/ - Ollama release notes —
https://github.com/ollama/ollama/releases - Kev (Jared Palmer) —
https://github.com/jaredpalmer/kev - lev (Interfaze) —
https://github.com/InterfazeAI/lev - decider (Mapika) —
https://github.com/Mapika/decider - Winnow-12B inference server —
https://github.com/EldanRing/winnow-inference - Ollaya —
https://ollaya.dev - Liquid AI decision models docs —
https://docs.liquid.ai/lfm/models/decision-models - Jev Decision Index (Hugging Face) —
https://huggingface.co/spaces/multimodalart/jev-decision-index - JevBench leaderboard —
https://benchmarkheaven.com/jev-models - Introducing System One Models and Jev (TypeSafe AI) —
https://typesafe.ai/blog/introducing-system-one-models-and-jev - System One concepts (TypeSafe docs) —
https://docs.typesafe.ai/concepts/system-one - Jev 1.13 jaggedness (TypeSafe docs) —
https://docs.typesafe.ai/model-jaggedness/jev-1.13 - The Register on Jev —
https://www.theregister.com/ai-and-ml/2026/09/16/typesafe-ai-debuts-model-for-machines-that-plays-doom/5296711 - InfoQ: TypeSafe AI releases Jev —
https://www.infoq.com/news/2026/10/typesafe-ai-jev-released/ - llama.cpp decision models (ggml-org) —
https://huggingface.co/blog/ggml-org/decision-models-in-llamacpp - Nimble (Bespoke Labs) —
https://github.com/bespokelabsai/nimble - Tev1 (Together AI) —
https://github.com/togethercomputer/tev1 - typesafe-sdk (Python) —
https://pypi.org/project/typesafe-sdk/ - Laya (Convai Innovations) —
https://huggingface.co/convaiinnovations/laya - Laya vs Jev: Open Weights and the Sovereignty Question (our September deep-dive, internal companion) —
https://kmail.at/blog/laya-vs-jev-open-weights-sovereignty
keep reading

Laya vs Jev: Open Weights and the Sovereignty Question
Two models defined the decision-model category in September 2026. Jev is a hosted frontier API; Laya is Apache 2.0 weights you can run on your own hardware. For a public body that difference decides whether the technology is admissible at all — measured against the EU four-level sovereignty framework.
ai · 18 min · 2026-09-27

Context Engineering: The Discipline That Ate Prompt Engineering
Prompt engineering asked what words; context engineering asks what configuration of context is most likely to produce the desired behavior — curated every turn, because the window is a budget. The full map: the 2025 coining, the measured evidence, the four levers, a 14-pattern playbook with code, and the honest counter-theses.
ai · 31 min · 2026-09-18

Classical ML is not dead: why gradient-boosted trees still beat deep learning on tabular data
Most of the data that decides whether a business works is a table. On tables, gradient-boosted trees still win — not from inertia, but from inductive bias. The published evidence, the mechanism, the operational reasons, and the exact line where deep learning takes over.
ai · 12 min · 2026-09-17