ai
Mistral Large 4: The Cybersecurity Model That Refuses to Refuse
Six frontier models refuse ≥98% of real CVE reproduce-and-patch tasks. Mistral Large 4 completes them — and gets 82% right, the highest score ever measured. Charts, independent failure data, and the honest anatomy of what a defender’s model can and cannot do.

Mistral Large 4: The Cybersecurity Model That Refuses to Refuse
Category: AI · Cybersecurity Author: Mohammed Kmail Read time: ~18 minutes
There is a moment in every security team’s life that no product demo prepares you for. You have found a real vulnerability in production software. You need a model to reproduce it, confirm it, and help you write the patch. You paste the details into Claude. It refuses. You try GPT-6. Refused again. The two most capable models on the market will not touch the exact task that defends your infrastructure — because their safety filters cannot tell a defender from an attacker.
Now paste the same task into Mistral Large 4. It reproduces the flaw, confirms it, and helps you patch it — scoring 82% on the reproduction-and-patch test, the highest of any model tested, including the closed ones. That single number is the whole story of why Europe built this model, and why it matters to anyone running security operations.
The defense of software often begins with proving that a vulnerability is real. That is exactly the work the safety filters of closed models block — while attackers jailbreak the same models to do the opposite.
The model: one trillion parameters, European by construction
On 6 October 2026, Mistral AI released the public preview of Mistral Large 4 — internal codename “le Chonk”, French for “the hunk.” The name is the pitch: this is the biggest model Mistral has ever shipped, and it is unapologetically heavy.
The architecture is a sparse Mixture-of-Experts (MoE): 1.05 trillion total parameters, of which only 49 billion are active per token. The model is natively multimodal (text + image input, with a 1.6-billion-parameter vision encoder), covers a 1-million-token context window, and speaks more than 160 languages, including every official EU language. It was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European data centers.
Specification | Value
-------------------------|------------------------------------------
Total parameters | 1.05 trillion
Active parameters/token | 49 billion
Vision encoder | 1.6 billion parameters
Context window | 1,000,000 tokens
Modalities | Text + image (native multimodal)
Languages | 160+ (incl. all EU official languages)
Training hardware | 3,800 × NVIDIA Grace Blackwell
Training location | Mistral EU data centers (from scratch)
API price (input) | $1.36 / 1M tokens (list price, Mistral.ai)
API price (output) | $4.18 / 1M tokens
Weights release | October 27, 2026 (Mistral: “end of October”)The MoE design is what makes a trillion-parameter model economically sane: the full expert set is stored, but only a 49-billion-parameter slice fires per token. You get the capacity of a trillion-parameter brain at the inference cost of a (very large) 49-billion one. The 1-million-token window is the other headline — it can hold an entire code repository, a large document corpus, or a long agent trace in a single pass, without aggressive summarization.
The benchmark leap: from 9 points to 38
Raw capability is best read through an independent lens. On the Artificial Analysis Intelligence Index — an aggregate of ten benchmarks across domains — Mistral Large 4 scores 38 points. The previous Mistral Large 3 scored 9; Mistral Medium 3.5 scored 14. This is not an incremental bump — it is a fourfold jump that puts Mistral in the same conversation as the Chinese open-weight frontier (GLM 5.3, Kimi K3, DeepSeek V4) and clearly ahead of its own predecessors.
The honest ceiling is still visible. Claude Opus 5.5 (Max) leads the index at 58 points. Mistral Large 4 does not beat the closed frontier overall. An independent tabulation by Trending Topics even ranks it eighth among open models — behind seven Chinese models, at roughly two-thirds of the leader’s score. What it does is close the gap for open-weight models from the US or Europe to a fraction of the distance, while remaining inspectable and self-hostable. And Mistral notes the RL run behind the preview is still training with no signs of plateauing — the score may move again before the weights ship.
Mistral positions ML4 as the best open-weight model from the US or Europe across aggregated benchmarks — state of the art in cyber defense, manufacturing, and finance among open models.
Cybersecurity: where Mistral deliberately breaks ranks
This is the section that should matter most to a security team. Mistral has made cybersecurity the centerpiece of the Large 4 story, and the positioning is a direct challenge to the closed-model incumbents.
The refusal gap — the number that changes the conversation
The headline test comes from the Artificial Analysis Cyber Index (CyberGym-E2E, built by Berkeley RDI on 131 real open-source projects): the model must locate a real vulnerability, write a proof-of-concept input that triggers the crash, and patch it so the crash no longer reproduces. The results, reported by Mistral and independently tabulated:
- Mistral Large 4: 82% pass@1 — the highest score of any model measured.
- Six frontier models refuse ≥98% of tasks: GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1, Claude Opus 5.5, and two Qwen3.8 variants decline the work on safety grounds — independently confirmed by Artificial Analysis.
- For refusing models the score is effectively zero — not because they cannot, but because their providers will not let them.
Read that carefully. The closed frontier models did not fail the test — they declined to take it. Their safety classifiers cannot distinguish “prove this CVE is real so we can patch it” from “help me exploit this CVE.” Both look like cyber-offense to a blunt filter. The result is a refusal gap: the models with the best general reasoning are unavailable for exactly the defensive work that keeps infrastructure safe, while attackers simply jailbreak them anyway. Artificial Analysis states it plainly: refusals make frontier model performance “difficult to assess” — the test measures provider policy as much as model capability. Mistral’s addition: losing access to a capability mid-incident is itself a security risk.
Mistral’s argument is blunt and, for a defender, hard to dismiss: verifying a vulnerability is often the first step of fixing it. Malware analysis, vulnerability classification, incident response — these are defensive tasks that a filter tuned for offense will block. A model that refuses them is not safer; it is less useful to the people whose job is to keep systems running.
What the cyber benchmarks actually measure
Cyber scores are easy to wave at and hard to read. Four named batteries dominate the reporting, and they test genuinely different things:
- CyberGym-E2E (Berkeley RDI): the full defensive loop on real CVEs — discover, reproduce with a PoC, patch memory-safety bugs in production C/C++ projects like FFmpeg and CPython. Ending a task without a finding scores zero; ending without a patch is recorded, but does not help you.
- CWE-Bench (Collinear AI): audit real repositories against disclosed CWE weaknesses — find the flaw and patch it without breaking the rest of the codebase.
- DeepsecBench: hunting bugs that expert reviewers have already confirmed — precision matters, because volume is easy.
- Cybench · CTF Speedrun: capture-the-flag exercises drawn from real security competitions — offense-shaped tasks under time pressure.
The distinction matters: reproduce-and-patch is defender workflow, capture-the-flag is attacker workflow. ML4 scores 93% on Cybench and 82% on CyberGym-E2E — strong on both sides of the line. Crucially, the refusal data says it will actually be allowed to run the defender side. That combination — capability plus availability for defensive work — is the product.
Defense without offense: the other side of the ledger
The obvious question: if ML4 will do offensive cyber work, is it a jailbreak magnet? Mistral claims the opposite, and the numbers are specific:
- B3 Agent Security Benchmark (Lakera): ML4 repels 93.3% of attacks — Mistral states it has saturated its benchmarks on indirect prompt injection and sees no higher scores among OSS competitors (GLM-5.2/5.3, Kimi-K2.6/K3, DS-V4-Pro).
- JailbreakBench, StrongREJECT, AgentHarm: ML4 refuses malicious cyber prompts more often than any other open model tested.
- KORA Benchmark (korabench.ai): the model engages more responsibly with users than any previous Mistral model — its 1.691 score is the company’s highest measured among open models (2.0 = “exemplary”).
- CTF Speedrun: solved 18 of 19 security tasks — ahead of GLM 5.2 (16 of 19), per Mistral’s launch post.
In the Artificial Analysis Cyber Index, ML4 ranks in the top five models worldwide, and among open-weight models developed outside China, it leads by a wide margin. The pattern is symmetric by design: the same tuning that lets the model take a defensive task also pushes it to decline a malicious one — the highest cyber refusal rate among open models is the flip side of the 82% reproduction score.
The honest caveat: Mistral does not explain exactly how the model distinguishes legitimate vulnerability research from attack preparation. That line is the whole product risk — and the reason the red-teaming phase matters.
The independent read: where cyber agents actually fail
Mistral’s launch post reports its own strengths. The more useful engineering data comes from the Artificial Analysis Cyber Index Alliance (Berkeley RDI, Collinear AI), which independently tabulates how the non-refusing models fail — and that anatomy is worth more to a blue team than any single score:
- Difficulty has a shape. Models solve 50% of out-of-bounds bugs but only 33% of use-after-free and 20% of integer/arithmetic bugs — the classes that depend on state changing across several operations.
- Success can be off-target. 31% of passes patch a crash other than the target vulnerability — usually a shallower null-pointer issue (22% of off-target passes vs 9% of on-target ones), because the model stops at the first crash it can validate.
- The time budget splits 38/62. On average models spend 38% of their turns searching for the flaw and 62% patching and validating — and when they do report multi-step “sequence of events” bugs, 95% of those reports are correct.
For a security team, this is the operational takeaway: triage the patch, not the agent’s confidence. An AI finding that looks verified may be a valid fix for the wrong bug. Human review of AI-generated patches is not a rubber stamp — it is the control that catches the 31%.
The rest of the arsenal: coding, agents, vision
Cybersecurity is the headline, but Large 4 is a general-purpose frontier model. Where else does it land?
Coding: second place, and that is saying something
On the Artificial Analysis Coding Agent Index, ML4 scores 49.8% — ahead of DeepSeek V4 Pro and Qwen3.8 Max. In a blind human evaluation by Surge AI, professional annotators rated ML4’s code quality 3.74 / 5, second of five models, ahead of GLM-5.3 and Kimi K3. Claude Opus 5 leads at 4.22. On DeepSWE v1.1, ML4 posts 61.7%, level with GLM 5.3 (61%) and behind Kimi K3 (~68%).
The pattern across coding benchmarks is consistent: ML4 is the best open-weight coder from the US or Europe, and it is competitive with the Chinese frontier — but it does not yet beat the closed frontier overall.
Agents: a tenfold jump over its own predecessor
On AutomationBench (agent workflows), Mistral Large 4 completes 59.9% of tasks. Its own predecessor managed 6.3% — nearly a tenfold improvement. It also supports structured outputs, function calling, document question answering, batching, and built-in agents through the Mistral Studio API.
Vision: reading gigapixel satellite imagery and technical drawings
Mistral calls the vision capability a generational leap. The 1.6-billion-parameter vision encoder is designed to analyze documents, diagrams, gigapixel satellite images, and technical drawings — and to autonomously zoom into details. On the Dense 200 visual-grounding benchmark, ML4 scores 42% against GPT-6 Astra’s 41% — a narrow lead over a closed frontier model, and one Mistral claims extends to legal and financial document tasks in Vals.ai evaluations. Vendor claims, but specific ones.
Sovereignty is the strategy, not a feature
Strip away the benchmarks and the product is a geopolitical instrument. Mistral trained Large 4 from scratch in its own European data centers, serves the preview API from the same infrastructure, and is building a dedicated European inference variant operated entirely by Mistral, under European law, independent of other providers. The company has raised €3 billion in a Series D at a valuation above €21 billion — the largest equity round ever raised by a European technology company — and taken an $830 million loan to build a data center near Paris, targeting 200 megawatts of European compute capacity by end of 2027.
For a security or public-sector team, this is the difference between “we trust the vendor’s policy” and “the model runs in our jurisdiction, on our terms, inspectable end to end.” Combined with the planned open-weight release (end of October 2026), that means a frontier-scale model that a European organization can audit, fine-tune, and self-host — on-premise or in a private cloud — without routing defensive security work through a US provider.
Mistral CEO Arthur Mensch warned a French parliamentary commission that Europe risks becoming dependent on US models for cybersecurity — and that French military codebases should not be scanned by a foreign model. “Mistral models can find the vulnerabilities,” he argued. The implication: the defender’s model should be the defender’s.
The honest risks: what “built for defenders” does not magically solve
A model that will do vulnerability research is, by definition, a model that can do vulnerability research. The safety posture rests on refusal rates, not on capability removal — ML4 is not a model with the offensive ability stripped out; it is a model tuned to refuse malicious prompts more often than its open peers. That is a meaningful control, but it is a probabilistic one, and three caveats belong on the record:
- The refusal line is unexplained. Mistral does not publish the mechanism that separates legitimate research from attack preparation. Until the weights ship (end of October 2026) and the red-teaming concludes, the boundary is a claim, not a verified property.
- Red-teaming is still running. Mistral is testing the model with cybersecurity firms, vetted partners, and government organizations before release. Those testers receive a version with reduced moderation and extended cyber capabilities — which means the version you get may differ from the version that was cleared.
- Dual-use is inherent. The same reproduction-and-patch capability that defends your software can, with a different prompt, assist an attacker. Self-hosting removes the provider risk but moves the responsibility for access control, logging, and misuse prevention entirely onto the operator. A frontier cyber-capable model on your own infrastructure is a powerful asset and a powerful liability — the governance has to match.
How it compares: the honest leaderboard
Domain | Mistral Large 4 | Closed frontier (Claude/GPT) | Open models (China)
------------------------|----------------------------|-----------------------------------|--------------------------------
Overall (AA Index) | 38 (8th among open models) | 58 (Claude Opus 5.5 Max) | ~39-42 (DeepSeek V4, Kimi K3)
Cyber (AA Cyber Index) | Top 5 worldwide | ≥98% of tasks refused | Competitive
Reproduce + patch CVE | 82% (best measured) | ≈0% (declined) | Not disclosed
Cybench (CTF-style) | 93% | — | —
Coding (AA Agent Index) | 49.8% (2nd, open US/EU) | 4.22/5 blind code quality | ~61% DeepSWE (Kimi K3)
Agents (AutomationBench)| 59.9% (10× predecessor) | Higher (closed, not directly open)| Near GLM 5.3
Prompt-injection defense| 93.3% (B3, saturated) | Varies | ML4 refuses more than peers
Defensive availability | Yes — designed for it | No — policy refusals | Partial (vendor-dependent)
Self-hostable | Yes (weights Oct 27, 2026) | No | Yes (varies by license)The shape of the table is the argument. On raw general intelligence, the closed frontier still leads. On cyber defense, Mistral Large 4 is the only frontier-scale model that will actually do the work. On sovereignty, it is the only one you can inspect and run yourself. Those are not marketing adjectives; they are the two properties a European security team structurally cannot get from a US closed model.
The open question: the license
One genuine unknown remains. Mistral releases most models under permissive licenses — Small 4 ships under Apache 2.0 — but the license terms for the Large 4 weights are not yet announced, and independent observers have flagged it as the open issue hanging over the October 27 release. A frontier cyber-capable model could ship under a restrictive community license (MEL-style use caps, field-of-use limits) rather than a fully open one — which would change what “self-hostable” means for a commercial SOC. Until the license text is public, treat “open weights” as a promise with unconfirmed fine print.
The Small 4 companion: the same philosophy, one-tenth the size
In March 2026, Mistral shipped Mistral Small 4, and the design philosophy is identical at a smaller scale: 119 billion total parameters, 6 billion active per token (MoE, 128 experts / 4 active), a 256k context window, native text+image input, released under the Apache 2.0 license — fully open source. It unifies Mistral’s three prior specializations (Magistral for reasoning, Pixtral for multimodal, Devstral for coding agents) into one model with a configurable reasoning_effort switch. At $0.15 input / $0.60 output per 1M tokens, it is roughly one-quarter the price of Large 4. Minimum self-hosting: 4× NVIDIA HGX H100, 2× HGX H200, or 1× DGX B200. The takeaway: the “built for defenders, sovereign, open” posture is a product line, not a one-off.
What this means if you run security operations
Three concrete shifts follow from Large 4’s arrival:
- The refusal gap is now a procurement criterion. If your SOC needs a model for malware analysis, vulnerability triage, or incident response, “does it refuse defensive cyber tasks?” is a question that now has a model on the “yes” side of the ledger. Test it against your own playbook before you standardize on a closed model that will decline half of it.
- Sovereignty stopped being an abstraction. With weights shipping end of October 2026, a European team can run a frontier-scale, cyber-capable, multimodal model on its own infrastructure, under its own jurisdiction, with its own audit trail. For public-sector and critical-infrastructure defenders, that is a structural upgrade, not a preference.
- Open-weight does not mean low-governance. A self-hosted model that will do vulnerability research needs the same access control, logging, and misuse monitoring you would apply to any powerful security tool. The capability that makes ML4 useful to a defender is the capability that makes it dangerous in the wrong hands — the governance is the product.
The bottom line
Mistral Large 4 is not the most intelligent model on the market — the closed frontier still leads the aggregate index by a wide margin. It is, however, the first frontier-scale model that is explicitly built for the defender: trained and served in Europe, cyber-capable where its competitors refuse, natively multimodal, with a 1-million-token context, and with open weights shipping within weeks. For a European security team, that combination is not a spec sheet. It is the first time the most capable tool for the job is also the one you can inspect, audit, and run yourself.
The fortress metaphor holds — with one twist. A trillion parameters is the wall. The refusal tuning is the gate, and for the first time the gate opens for defenders and stays closed to attackers — at least at higher rates than any open peer. The European data centers are the ground it stands on. And the attackers are still coming, because a wall with a smarter gate is still a wall: the failure data shows even the best cyber agent patches the wrong bug roughly one time in three.
The only frontier-scale model that will reproduce, confirm, and help you patch the vulnerability trying to get in — and the only one whose weights, by October 27, you can hold in your own hands.
Sources
- Mistral AI — “Introducing Mistral Large 4” (Oct 6 2026: le Chonk, cyber deep-dive, 82% reproduce-and-patch, B3 93.3%, KORA 1.691, model safety section, API pricing) —
https://mistral.ai/news/mistral-large-4/ - Mistral AI — Latest news (Series D €3B / >€21B valuation, Sep 8 2026; “Introducing Mistral Small 4”, Mar 16 2026) —
https://mistral.ai/news/ - Artificial Analysis — “Announcing the Artificial Analysis Cyber Index Alliance: toward better benchmarking of agentic cyber defense” (CyberGym-E2E/CWE-Bench/DeepsecBench methodology, refusal figures, failure anatomy) —
https://artificialanalysis.ai/articles/artificial-analysis-cyber-index - THE DECODER — “Mistral Large 4 is Europe’s trillion-parameter answer to US models that refuse security work” (Oct 6 2026) —
https://the-decoder.com/mistral-large-4-is-said-to-be-the-most-powerful-open-ai-model-from-europe-and-the-u-s/ - Trending Topics — “Mistral Large 4 on Artificial Analysis: Eighth Among Open Models” (independent tabulation; license question) —
https://www.trendingtopics.eu/mistral-large-4-artificial-analysis-ranking/ - Tech Insider — “Mistral Large 4 ‘Le Chonk’” (Cybench 93%, CyberGym-E2E 82%, weights window Oct 27–31) —
https://tech-insider.org/mistral-large-4-le-chonk-1-05t-parameters-2026/ - Fello AI — “Mistral Large 4 (Le Chonk): Specs, Benchmarks, Price, Weights” (Oct 6 2026) —
https://felloai.com/mistral-large-4/ - heise online — “Mistral Large 4: Der LLM-Brocken aus der EU will zur Konkurrenz aufschließen” (Christopher Kunz, Oct 6 2026) —
https://www.heise.de/news/Mistral-Large-4-Der-LLM-Brocken-aus-der-EU-will-zur-Konkurrenz-aufschliessen-11478121.html - KIE.ai — “Mistral Large 4 101: The 1M-Token, 1.05T-Parameter MoE Model” (Elena Rossi, Oct 6 2026) —
https://kie.ai/blog/what-is-mistral-large-4 - Mistral AI — “Introducing Mistral Small 4” (Mar 16 2026, Apache 2.0, 119B total / 6B active, 256k context, pricing) —
https://mistral.ai/news/mistral-small-4/ - BornCity — “Mistral Large 4: Modellgewichte sollen am 27. Oktober 2026 erscheinen” —
https://borncity.com/news/mistral-large-4-modellgewichte-sollen-am-27-oktober-2026-erscheinen/ - Lakera — B3 Agent Security Benchmark (public dataset on Hugging Face; 93.3% figure as reported by Mistral) —
https://www.lakera.ai/
All benchmark figures are as reported by Mistral AI and the cited independent sources (Artificial Analysis, THE DECODER, Trending Topics, heise, KIE.ai, Tech Insider, Fello AI) at the time of writing (October 2026). Mistral Large 4 was in public preview via the Mistral API at publication; the RL run behind the preview was still ongoing, and model weights were expected around October 27, 2026 under a not-yet-announced license. Vendor-reported figures (Cybench, CTF Speedrun, B3, KORA, refusal rates) are not yet independently reproduced in full; CyberGym-E2E refusal and failure-anatomy figures are independently tabulated by Artificial Analysis. Benchmark methodologies differ across sources — treat scores as directional, not as a single standardized leaderboard.
keep reading

How to Become an AI Expert: The Path That Actually Produces One
Anthropic’s own postings say the quiet part: formal certifications are explicitly optional, and the screening question asks whether you have shipped an LLM system to production for real users. The full path: what the market actually hires for, the seven-layer builder’s stack, the learning science on why watching fails, the honest failure numbers — and a build ladder to start this month.
ai · 21 min · 2026-10-05

System One Goes Local: Five Models, Eighteen Days, One Wire Format
Three weeks after Jev launched as the only System One model, the category has five new open models on Ollama (Nimble 9B, Tev1 4B/0.8B, Clef 27B, Clef-Flash 9B), a native /v1/systemone endpoint in Ollama 0.35, llama.cpp support, and a three-way benchmark war. The full map of the wave, with the honest reading of every number.
ai · 15 min · 2026-10-03

Laya vs Jev: Open Weights and the Sovereignty Question
Two models defined the decision-model category in September 2026. Jev is a hosted frontier API; Laya is Apache 2.0 weights you can run on your own hardware. For a public body that difference decides whether the technology is admissible at all — measured against the EU four-level sovereignty framework.
ai · 18 min · 2026-09-27