Mistral put a public preview of Large 4 on its blog on October 6. The company calls the model ML4 and, unofficially, Le Chonk. It is a natively multimodal mixture-of-experts system with 1 trillion parameters and 49 billion active parameters. The preview API is live in Mistral Studio. Weights are due at the end of the month.
Training ran from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own European data centers. The preview is served on that same stack. The post says a large share of the training data covered more than 160 languages, including every official language of the European Union. The model card lists a 1 million-token context, a 1.6 billion-parameter vision encoder, and preview prices of $0.68 per million input tokens and $2.09 per million output tokens.
On the numbers Mistral published, Large 4 scored 61.7 percent on DeepSWE v1.1 and 28.3 percent on Terminal-Bench 4. A combined coding-agent score of 49.8 percent is placed ahead of DeepSeek V4 Pro and Qwen3.8 Max. On Cybench it solved 93 percent of 40 competition exercises. On one Artificial Analysis cyber test — reproduce a real vulnerability, then patch it — the company reports 82 percent, and says Claude Opus 5.5 and GPT-6 Astra score near zero on that same task because they refuse it.
Key takeaways
- Large 4 preview opened October 6: 1 trillion parameters, 49 billion active, weights due by month’s end.
- Trained on 3,800 Grace Blackwell GPUs in Mistral’s European data centers; 1 million-token context on the model card.
- Company figures: 61.7 percent DeepSWE v1.1, 93 percent Cybench, 82 percent on an Artificial Analysis cyber reproduction test.
What does Mistral Large 4's preview actually ship?
A hosted API, not a download. Studio users can call the preview now. The weights, the part that would let a lab run the model without Mistral in the request path, are still scheduled rather than posted. Until that drop, Mistral says it is red-teaming the system with cybersecurity teams, vetted partners, and state authorities, on a version with reduced moderation and wider cyber tools.
The refusal point is the product argument, not a side note. Mistral’s post says closed models can block the first step of defensive work — proving a flaw is real — while the same models get jailbroken for offensive use. Large 4 is pitched as the open-weight alternative that still refuses more malicious cyber prompts, on JailbreakBench, StrongREJECT, and AgentHarm, than other open systems the company compared. That claim is Mistral’s own measurement. Independent ranking is thinner: the announcement does not publish a full leaderboard placement.
For creators wiring model calls into a production stack, the relevant split is price and control. The card’s $0.68 / $2.09 rates sit under several closed frontier lists, and the end-of-month weights are the piece that changes hosting. A longer look at how those choices sit in a working setup is in How AI is Reshaping the Content Creator Stack in 2026.
What to watch next
The weight release, which Mistral has dated only as the end of October, and whether the DeepSWE and Cybench figures hold once outside labs can rerun them on the open checkpoint. A second item is the cyber-access list: who gets the reduced-moderation preview before the weights are public.
