A public body cannot audit a decision it did not run. That single sentence explains why the most consequential thing about the new class of decision models is not their accuracy — it is whether their weights fit on hardware you control.
In September 2026, two models defined a new category at almost the same moment. TypeSafe AI emerged from stealth with $40 million in seed funding led by DCVC and released Jev, the first publicly available System One model — a frontier-class AI that generates no text at all, returning typed decisions and calibrated probabilities instead. Weeks earlier, Convai Innovations had already published Laya: the same architectural idea, the same three decision primitives, the same training method family — but with Apache 2.0 weights anyone can download and run.
They are close cousins on paper and near-opposites in procurement. Jev is a hosted API. Laya is a file you hold. For a ministry, a statistical office, or a market-surveillance authority, that difference decides whether the technology is admissible at all.
Every mainstream large language model — GPT, Claude, Gemini, the open Llama family — works the same way: it generates text one token at a time, and your software then has to parse meaning back out of that string. The model writes the answer, and you hope your parser agrees with it.
A System One model inverts the pipeline. You send a block of program state plus a set of typed questions, and the model evaluates all of them in a single parallel pass, returning structured values with probabilities. There is nothing to parse, because nothing was generated. The output is already in the shape your code branches on.
The category name comes from Daniel Kahneman’s Thinking, Fast and Slow — System 1 being fast, intuitive judgment; System 2 being slow deliberation. The bet is that most software decisions need the first, not the second.
Both Laya and Jev expose the same three question primitives:
choice — pick one option from a set you name. Returns a label plus a probability distribution over all options.score — place the input on an ordered scale you define. Returns a value on that scale.noul — estimate the probability that a proposition is true. Returns a single calibrated probability.That is the whole interface. It looks almost insultingly small next to a chat model’s generality, and that is precisely the point: constraint is what makes the output dependable inside a system with latency guarantees.
Laya is worth understanding at the level of mechanics, because its architecture explains both its speed and its failure modes.
The backbone. Laya is not a small language model. It is a bidirectional encoder — ModernBERT-large, fully fine-tuned — with a purpose-built decision head trained from scratch on top: two transformer layers, an option-marker scorer, and an act/escalate head. The English checkpoint totals 421M parameters. The multilingual checkpoint swaps in mmBERT-base, 22 layers and a 256k vocabulary, at 322M parameters.
The important structural fact is the option-marker scorer. Each candidate option is scored at its own mask token, and the scores are softmaxed across that question’s options. Because the answer space is supplied at request time rather than baked into the weights, a new question schema needs no retraining. You can invent a new taxonomy on Tuesday and use it on Tuesday.
The forward pass. A non-autoregressive model produces its entire answer set in one pass, not one token at a time. There is no reasoning trace and no chain of thought — nothing to sample sequentially, so nothing to go off the rails mid-generation. On a single Tesla T4, Laya answers one question in about 33 ms, and the cost per question collapses under batching: 7.2 ms per question at a batch of ten, reaching 103 to 332 questions per second on one GPU.
The Router. Laya ships a language and script router that detects the script in under half a millisecond in pure Python before the forward pass, then dispatches to the checkpoint that can actually read the input. This matters more than it sounds. Laya’s own documentation records that the English checkpoint collapses on non-Latin scripts — Khmer scores 0.000 accuracy at 0.952 confidence. The model stays confident while being wrong, which means confidence gating alone cannot save you. Routing is the fix, and it is the reason the multilingual checkpoint exists as a separate artefact rather than one blended model.
The training method. Both models use a method called RLCD — Reinforcement Learning for Calibrated Decisions. The policy reports a distribution, exploration adds zero-mean Gaussian noise to the logits, and the reward is a strictly proper scoring rule: log and spherical for categorical questions, a ranked probability score for ordinal ones. Laya’s implementation is REINFORCE with a group-mean baseline, GRPO-style. The property that matters is mathematical rather than empirical: under a strictly proper scoring rule, expected reward is maximised only by reporting honest probabilities. The model is not asked nicely to be calibrated; it is scored in a way that makes honesty the optimal strategy.
Here is the comparison, with the caveat that belongs on it stated first: Laya’s benchmark table is self-published. The repository states plainly that Jev figures were never measured locally — there was no TypeSafe API access — so sample sizes and prompts differ between columns. Treat the accuracy deltas as indicative, not as an independent verdict.
POST /v1/systemone, Jev-compatible · Jev (TypeSafe AI): POST /v1/systemoneThe headline claim — that an open 421M-parameter encoder matches or beats a funded frontier lab’s flagship on typed decisions — is a real result if the harness was fair, and the repository does document its methodology: byte-identical questions per model, fixed seeds, 400 cases per task, full result JSON committed to the repository.
The honesty cuts both ways, and this is where an expert differs from a press release. Jev wins decisively on high-cardinality choice. On Banking77, Laya scores 0.425 against Jev’s 0.870. The mechanism is documented and mundane: Laya’s option-marker head allocates a fixed token budget across options, so at default settings a 77-option question gives each label roughly three to four tokens, and accuracy falls off a cliff. Laya’s own mitigation is to raise head_max_len and max_len, or to restructure the decision as a coarse-to-fine two-step choice.
The practical reading: Laya is not a strict superset of Jev. Below roughly twenty options they are comparable or Laya leads; above that, Jev’s architecture handles a wide option space that Laya’s needs help with. Anyone claiming an unqualified win is selling something.
One more convergence deserves attention. Laya’s laya-serve exposes the Router over the same POST /v1/systemone wire protocol as TypeSafe’s hosted API. An existing TypeSafe client changes its base URL and talks to self-hosted weights instead. The switching cost between the open and closed implementations of this category is, quite literally, one line of configuration.
The comparison that matters for an organisation is not Laya versus Jev. It is the decision layer versus the chat model you already have.
Output shape. An LLM returns a string; the value is latent in the text and must be extracted by a parser, validated against a schema, and defended against the case where the model returns something adjacent to what you asked for. A System One model returns the typed value directly. There is no parsing step to fail, because the possible outputs were constrained in advance.
Hallucination. This is where the marketing needs a scalpel rather than a headline. TypeSafe claims Jev "can’t hallucinate", and The Register’s coverage pushed back fairly: the claim is not a like-for-like comparison, because Jev does not produce natural language, and type-safety "does not preclude the possibility of being incorrect". Both things are true. A decision model cannot fabricate a citation, invent a statute, or confabulate a tool call, because it never writes prose. It can absolutely choose the wrong label from a correct option list. Type-safety eliminates the entire failure mode of malformed output; it does not touch the failure mode of wrong judgment.
Latency. A frontier LLM’s end-to-end response time runs into seconds, often tens of seconds, for structured work. That is perfectly fine for a human waiting on a chat reply and fatal for anything inside a live event loop. The decision layer operates in tens of milliseconds, which is what makes continuous classification, inline filtering, and real-time routing economically possible rather than merely desirable.
Cost. Self-hosted Laya has no per-decision price; the cost is hardware you already own. Jev charges $0.042 per million input tokens with output free. TypeSafe’s own framing is that a decision costs on the order of four hundred-thousandths of a dollar, which is what makes scoring an entire table — fifty million rows for roughly twenty dollars — a normal operation instead of a budget line.
Calibration. Language models are overconfident even when explicitly asked for a confidence estimate. A model that is right 95% of the time but cannot flag the other 5% cannot be given unattended authority. Calibrated probabilities are the mechanism that converts "AI that usually works" into "AI that can run a million times without a human", because confidence gives you a threshold: auto-act above it, route the uncertain remainder to a person or a reasoning model.
What each is for. The honest division of labour is a split stack. The decision layer classifies, scores, routes, and guardrails at volume and speed. The chat model keeps the genuinely open-ended work — drafting, explaining, reasoning through novel problems, writing code. Neither replaces the other; they occupy different tiers of the same architecture.
Here the argument stops being about model quality and becomes about law, procurement, and liability. A public body is not a startup optimising for time-to-market. It is accountable for the decisions it delegates, and it must be able to demonstrate that accountability after the fact.
The regulatory direction of travel is explicit. The European Union’s Cloud and AI Development Act establishes a single sovereignty framework with four assurance levels that public bodies are expected to use according to their own risk assessment:
The Commission’s Cloud Sovereignty Framework names the same test under SOV-3, Data and AI Sovereignty: the extent to which AI models and data pipelines are developed, trained, hosted, and governed under EU control, minimising dependency on non-EU technology stacks. It also requires strict confinement of storage and processing to European jurisdictions, with no fallback to third countries.
Read that against the two models. A hosted API is a service relationship. You are a customer of an inference endpoint operating under someone else’s jurisdiction, and your data residency, your availability, and your continued access are contractual rather than physical facts. Open weights are a different category of thing entirely. The model is an artefact on your disk. There is no external dependency at inference time, because there is no external party at inference time.
The EU has already legislated that instinct. The EU Open Source Strategy makes technological sovereignty its first objective, names public administrations as anchor users and contributors, and commits to procurement guidance and open-source-friendly tendering. It explicitly aims to reduce dependence on non-EU technologies and to increase control over critical digital infrastructure. For the EU Digital Identity Wallet, the European Digital Identity Regulation goes further and makes open source the legal default for application software components.
The strategy is not without its critics, and the criticism is useful context. The Centre for European Policy notes that the Open Source Maintenance Instrument is estimated at roughly €2 billion over seven years — thin, it argues, against €264 billion in annual proprietary IT spending, and that funding should be tied to maintenance lifecycle obligations rather than project launches. The lesson for a public body is not to wait for the policy to mature. It is that the strategic direction is settled, and the procurement decisions taken now determine who is positioned when the framework hardens.
Four kinds of sovereignty, one architectural choice. Digital sovereignty is commonly decomposed into governance, technical, operational, and data sovereignty. A decision model touches all four, and self-hosting is what makes them achievable rather than aspirational:
The auditability problem, stated plainly. System One models deliberately return no rationale. You get a label and a probability, not an explanation. For a market-surveillance authority or a statistical office operating under the EU AI Act, "the model said so" is not a defensible record. This is a genuine architectural limitation, not a teething problem, and it has a real consequence: the decision layer is not a substitute for human accountability, it is a component inside a system that must supply that accountability itself.
Which is why the pattern that works is confidence-thresholded escalation. The decision layer handles the high-volume, high-confidence majority; everything below threshold is routed to a slower model that can explain itself, or to a human reviewer. The audit trail comes from the application around the model — the state that went in, the decision that came out, the probability attached to it, and the rule that determined what happened next. If that sounds like more work than calling an API, it is. It is also the difference between automation you can defend to an auditor and automation you cannot.
And the practical argument that decides it. A public body that runs open weights keeps its capability through a vendor’s acquisition, a price change, a licence revision, an export-control decision, or a geopolitical rupture. Those are not hypothetical risks; they are the stated motivation for the entire EU sovereignty package. A 421M-parameter model that fits in 808 MB and runs on commodity hardware is procurable in a way that a dependency on a foreign startup’s endpoint is not — because there is nothing to procure after the first download.
The honest scope condition is narrow and large at the same time: high-volume, repeated decisions over a shared state, where the valid answers are known in advance. That describes an enormous share of what software in a public body actually does.
Case routing and triage. Incoming correspondence, applications, and citizen requests get classified by subject and urgency, and routed to the correct unit. The option space is a fixed organisational taxonomy, which is precisely what a choice question is for.
Document relevance and screening. Given a case file and a set of criteria, decide whether each document is relevant. This is noul at volume — the kind of map-reduce over a large corpus that was uneconomical when every row required a multi-second LLM call.
Guardrails and verification. Score model inputs and outputs for policy violations, prompt injection, or jailbreak attempts before they reach a downstream system. This is where the latency matters most, because a guardrail that takes three seconds is not a guardrail.
Anomaly and risk scoring. Assign a calibrated risk score to a submission, a transaction, or a declaration, and set a review threshold. Because the score is a probability rather than a label, the threshold becomes a tunable policy instrument instead of a hard-coded rule.
Model routing. Decide whether a request needs the expensive frontier model or can be answered by something smaller. This is the meta-application: use a cheap decision to avoid an expensive generation, and it pays for itself immediately in any system already paying per token.
Formal-complaint and escalation detection. Detect whether a communication constitutes a formal complaint, a legal challenge, or a request that starts a statutory clock, and escalate accordingly. The consequence of missing one is severe enough that a calibrated probability plus a human threshold is a better design than regex.
Structured extraction over document sets. Turn a pile of unstructured text into rows — a feature, a flag, a category per document — that downstream analytics can query. This is the map-reduce pattern: cheap per-decision cost is what makes whole-corpus analysis feasible.
Real-time and interactive applications. Anything that must respond inside a live event loop: adaptive interfaces, simulation, streaming content checks. At 33 ms, the decision is faster than the user’s perception of the interaction.
A technology assessment that omits the failure modes is not an assessment. These are Laya’s own documented weaknesses, taken from its model card and benchmarks rather than inferred:
score is the weakest primitive. Measured at 0.372 on SST-5. If your decision is genuinely about degree rather than category, expect the least reliable behaviour from the primitive you need most.noul can be distracted by its own option labels. The documentation records answers following the false: / true: labels rather than the state, and recommends rephrasing as a two-option choice with neutral keys when answers look stuck.The boundary this draws is useful rather than disqualifying. Use the decision layer where the answers are bounded and the volume justifies it. Keep deterministic logic in code. Keep explanation and open-ended reasoning with the models built for it. And validate calibration on your own distribution before anything runs unattended.
The decision layer is the missing infrastructure that makes AI reliable enough to be boring — and 2026 is the year it became a choice between two philosophies rather than one product.
Jev is the more polished frontier implementation. If your workload is high-cardinality, if you are commercial, and if a hosted API is acceptable, it is a serious piece of engineering from a team that knows this problem deeply.
Laya is the one you can hold. It matches or beats Jev on typed decisions and calibration in its own published comparison, it loses clearly on wide option spaces, and it runs on hardware you own, under a licence that lets you inspect, modify, and redeploy it. For a public body, that last property is not a preference. Under the sovereignty framework Europe is actively legislating, it is the requirement that everything else is measured against.
If you are in the public sector and want to evaluate this category, the sequence is concrete. Pick one bounded, high-volume decision your organisation already makes by hand. Build an evaluation set from six months of real cases with known outcomes. Run Laya in shadow mode against your existing process, logging every decision, probability, and disagreement without letting it act. Recalibrate the temperatures on your own data. Only then set a confidence threshold and let the high-confidence slice run automatically, with the remainder escalating to a human.
That is slower than calling an API. It is also the only version of this that an auditor will accept — and once it is running, the marginal cost of the millionth decision is the same as the first.
Sources for every claim in this article, for readers who want to verify rather than take on faith.
Laya — primary sources
https://github.com/NandhaKishorM/layahttps://huggingface.co/convaiinnovations/layahttps://github.com/NandhaKishorM/laya/blob/main/BENCHMARKS.mdhttps://nandhakishorm.github.io/laya/https://nandhakishorm.github.io/laya/staged-adoption/https://pypi.org/project/laya/https://huggingface.co/spaces/convaiinnovations/laya-demoJev / TypeSafe — primary sources
https://docs.typesafe.ai/concepts/system-onehttps://typesafe.ai/blog/introducing-system-one-models-and-jevhttps://www.theregister.com/ai-and-ml/2026/09/16/typesafe-ai-debuts-model-for-machines-that-plays-doom/5296711EU sovereignty framework — primary sources
https://digital-strategy.ec.europa.eu/en/policies/cloud-and-ai-development-acthttps://commission.europa.eu/document/download/09579818-64a6-4dd5-9577-446ab6219113_enhttps://digital-strategy.ec.europa.eu/en/policies/open-source-strategyhttps://data.consilium.europa.eu/doc/document/ST-10220-2026-INIT/en/pdfhttps://www.cep.eu/eu-topics/details/eu-tech-sovereignty-package.htmlBackground concepts
https://www.bundesdruckerei.de/en/innovation-hub/digital-sovereignty-what-is-it