Mistral launched a public preview of Mistral Large 4 on October 6, 2026: a 1-trillion-parameter, natively multimodal model with 49 billion parameters active per token. It is Mistral’s largest model, trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in its own European data centres. Mistral says the open weights will be released by the end of October. The preview API costs 1.36 dollars per million input tokens and 4.18 per million output. The company’s own nickname for it is "le Chonk."
How good is it?
Strong for an open model, behind the closed frontier. On coding, Mistral reports 61.7% on DeepSWE v1.1, 59.4% on SWE-Atlas-QnA and 28.3% on Terminal-Bench 4.0, for a combined Coding Agent Index of 49.8% that puts it ahead of DeepSeek V4 Pro and Qwen3.8-Max. On AutomationBench, 657 business workflows across tools like Gmail and Salesforce, it scores 59.9%.
For context, Claude Opus 5.5 scores 66.4% on Terminal-Bench 4.0 against Large 4’s 28.3%. In a blind human evaluation of coding quality, Large 4 ranked second of five behind Claude Opus 5. Its preview scored 38 on the Artificial Analysis Intelligence Index, the best Western open-weight result but still behind the leading Chinese open models.
What is the cybersecurity claim?
This is the part worth reading carefully. Mistral says Large 4 ranks in the top five on the Artificial Analysis Cyber Index and scores 82% on a test that asks a model to reproduce a real vulnerability in open-source software and then patch it, the highest of any model. It solves 93% of the 40 Cybench challenges.
Mistral also points out that Claude Opus 5.5 and GPT-6 Astra score near zero on that same reproduce-and-patch test because they refuse the task. That is a genuine point about defenders being blocked mid-incident. It is also exactly the capability the closed labs have deliberately restricted, through programmes like Google’s Fairwind and Anthropic’s cyber verification tiers. Shipping it as open weights means anyone can run it without a vetting step, which cuts both ways.
Mistral says it is red-teaming the model with cybersecurity firms, vetted partners and state authorities before the weights go out, and that Large 4 refuses malicious cyber prompts more often than other open models on JailbreakBench, StrongREJECT and AgentHarm.
Why does "trained in Europe" matter?
Sovereignty. Mistral trained and serves the model on its own infrastructure and will offer a European deployment it runs end-to-end under European law, independent of US cloud providers. For European governments and regulated firms that cannot send data to American services, that is the selling point more than any benchmark. The model is the first output of Mistral’s 3 billion euro Series D, which it calls the largest equity round ever raised by a European tech company.
Should you use it?
If you need open weights, European data residency or security research without refusals, it is now the strongest Western option. If you need the best coding agent available, the closed frontier models are still well ahead on Terminal-Bench. And the usual caveat applies: most numbers above are Mistral’s own or privately run evaluations, and the reinforcement learning run is still in progress, so the released weights may score differently. It continues the open-model squeeze on pricing we covered with AT&T moving traffic to open models.




