For the past two years, the best models you could download and run yourself have come from China. DeepSeek, Qwen, Kimi and GLM set the pace, while Western labs kept their strongest work behind an API. On 6 October 2026, the French company Mistral put up a public preview of Mistral Large 4 — unofficially ML4, and officially, in Mistral's own words, le Chonk — and claimed the gap is closing. The weights are promised by the end of the month.
- 1 trillion parameters, 49 billion active. A natively multimodal, granular mixture-of-experts model with a 1.6-billion-parameter vision encoder and a 1 million token context window.
- Open weights, but not yet. You can use the preview API today on Mistral Studio; Mistral says it will release the weights by the end of October 2026.
- Trained entirely in Europe on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own datacentres, on data spanning more than 160 languages, including every official language of the European Union.
- It refuses to refuse security work. On a test that asks a model to reproduce a real vulnerability and then patch it, ML4 scores 82% — the highest of any model — while Claude Opus 5.5 and GPT-6 Astra score near zero because they decline the task.
- The careful claim: Mistral says ML4 is "competitive with the strongest open-source models globally" while "significantly outperforming any open-weight model developed in the US or Europe". Read that second phrase closely.
What Mistral Large 4 actually is
ML4 is a mixture-of-experts model. That means it is not one enormous network that runs end to end for every word; it is a large collection of smaller expert sub-networks with a router that picks a handful for each token. Mistral's announcement gives 1 trillion total parameters with 49 billion active, and its model page lists 1.05 trillion total and 52 billion active — the two pages differ slightly, which is worth knowing if you are quoting the figure.
Either way, the ratio is the point. Only around 5% of the model does any work for a given token, which is what makes a trillion-parameter model affordable to serve at all. If you want to understand the mechanism properly, our lesson on why mixture of experts covers it, and you can route tokens to experts yourself in the Transformer simulator.
It is natively multimodal, with a 1.6-billion-parameter vision encoder built in rather than bolted on, and it takes up to 1 million tokens of context.
What it costs
Mistral's model page lists ML4 at $1.36 per million input tokens, $0.14 per million cached input tokens and $4.18 per million output tokens. At the time of writing those are discounted to $0.68, $0.07 and $2.09 respectively — exactly half. Check the current rate before you budget anything, because the discount is promotional.
Where it is strong
Mistral published a wide set of benchmark numbers. These are the company's own figures, and nobody has independently reproduced them yet, so treat them as claims rather than settled results.
| Area | Benchmark | Score |
|---|---|---|
| Cybersecurity | Vulnerability reproduce-and-patch (Artificial Analysis Cyber Index) | 82% — highest of any model |
| Cybersecurity | Cybench (40 security-competition challenges) | 93% |
| Coding | DeepSWE v1.1 | 61.7% |
| Coding | SWE-Atlas-QnA | 59.4% |
| Coding | Terminal-Bench 4 | 28.3% |
| Coding | Coding Agent Index (combined) | 49.8%, ahead of DeepSeek V4 Pro 0813 and Qwen3.8 Max |
| Agents | AutomationBench (657 business workflows) | 59.9% |
| Agents | AA-Briefcase (long-horizon knowledge work) | 1,393 Elo |
| Vision | Dense 200 visual grounding | 42%, against GPT-6-Astra's 41% |
| Safety | Lakera B3 attack resistance | resists 93.3% of attacks |
In a blind human evaluation run with Surge AI, professional annotators rated coding output on a 1–5 scale without knowing which model produced it. ML4 Preview came second of five at 3.74, ahead of GLM-5.3 (3.60), Kimi K3 (3.59) and GLM-5.2 (3.40), and behind Claude Opus 5 (4.22). A second-place finish, honestly reported against a closed model that beat it, is more persuasive than most benchmark tables.
The genuinely interesting part: refusals
The cybersecurity results are the most consequential thing in this release, and not because of the number.
Mistral points out that several leading closed models, including Claude Opus 5.5 and GPT-6 Astra, score near zero on the vulnerability reproduce-and-patch test — not because they cannot do it, but because they refuse. Their safety filters treat "reproduce this vulnerability" as an attack.
Mistral's argument is that defending software usually starts with proving a flaw is real, so a filter that blocks exploit work also blocks the defender. It adds that losing access to a capability in the middle of an incident is itself a security risk, and that attackers are already jailbreaking those same closed models for offensive work. Open weights that a security team can run on its own hardware, under its own policy, side-step the problem.
That is a real argument, and it is also a commercial one — Mistral sells to exactly the organisations who find refusals expensive. Both things are true at once. Worth noting too that Mistral says it is red-teaming ML4 with cybersecurity leaders, vetted partners and state authorities, who get access to a version with reduced moderation, before the weights go out.
Mistral also reports that despite those capabilities, ML4's average refusal rate on malicious cyber prompts from JailbreakBench, StrongREJECT and AgentHarm is higher than every other open-source model it compared against, and that it scores 1.691 out of a maximum 2 on the KORA benchmark for responsible engagement.
"Forged in Europe"
ML4 was trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacentres, and the preview is served from that same infrastructure. Mistral is explicit that this is about sovereignty: a European deployment it operates end to end, independently of other digital service providers, under European law.
For Indian readers the sovereignty argument is a familiar one, and the multilingual detail is the more interesting number: a significant share of the training data spanned more than 160 languages.
This is the first milestone from Mistral's €3 billion Series D, which the company says is the largest equity round ever raised by a European technology company.
How it was trained
Mistral is unusually open about its reinforcement learning setup. Its RL library combines very different tasks in a single run — chat, scientific problem solving, safety alignment, factuality, long-horizon tool use — sharing code sandboxes, web search and external APIs. An autoscaling fleet generates tens of thousands of rollouts in parallel while training proceeds asynchronously.
At roughly 3,000 GPUs, Mistral says a single training run produces about 33 billion tokens per day, of which around 16 billion are trainable completion tokens after filtering. The company adds that the RL run behind this preview is still going and shows no sign of saturating, so it expects the model to improve further before the weights are published.
What to watch
The weights are the whole story, and they are not out yet. Until they are, every number above is a claim from the company that made the model. Three things worth watching:
- Whether the weights actually land in October, and under which licence — "open-weight" is not the same as open source, and Mistral has not named the licence.
- Whether independent evaluations reproduce the cyber results. The Artificial Analysis Cyber Index is public, so this is checkable.
- Whether the refusal argument holds up. A model that will do security work for defenders will also do it for attackers, and Mistral's own mitigation is a higher refusal rate on explicitly malicious prompts. That balance is the thing to judge once anyone can run it.
If you want the background on why a trillion-parameter model can run at all, start with why mixture of experts and why low precision, then look at anatomy of a frontier model.
Resources
- Mistral: Mistral Large 4 announcement — the full launch post, with the benchmark charts
- Mistral model documentation — parameters, context length and current pricing
- Mistral Studio — the preview API
- Mistral on Hugging Face — where the open weights are expected at the end of October 2026
- Artificial Analysis — the independent index behind the cyber scores