Mistral Large 4: a trillion-parameter model trained entirely on European GPUs
61.7% on DeepSWE, a visual-grounding win over GPT-6-Astra, and weights due by month's end -- Mistral's answer to the frontier-model race is also a bet that training capacity doesn't have to come from a US hyperscaler.
Mistral AI released Large 4 on October 6, a natively multimodal mixture-of-experts model with 1 trillion total parameters and 49 billion active parameters, trained on 3,800 Nvidia Grace Blackwell GPUs inside Mistral's own European datacenters. The training run covered more than 160 languages, including every official EU language -- a detail Mistral is pointedly making part of the pitch, not an afterthought. Public preview is live now through Mistral Studio, priced at $1.36 per million input tokens and $4.18 per million output tokens; open weights are planned for "end of this month."
The numbers Mistral chose to lead with
On Mistral's own benchmark table, Large 4 scores 61.7% on DeepSWE v1.1, 93% on 40 Cybench challenges, and 82% on the AA Cyber Index's vulnerability-reproduction test -- a cluster of agentic-coding and security benchmarks, not general knowledge tests, which is itself a signal of who Mistral is building this model to compete for. The one head-to-head comparison in the release names a specific rival directly: Large 4 beats "GPT-6-Astra" 42% to 41% on a visual-grounding benchmark (Dense 200) -- a single-point margin Mistral is still choosing to publish and name.
Why the GPU provenance is the actual story
Every frontier-model release competes on benchmarks; fewer compete on where the compute came from. Training a trillion-parameter model entirely inside European infrastructure, on hardware Mistral operates rather than rents from a US hyperscaler, is a sovereignty argument aimed squarely at the enterprise and government buyers for whom "which country's cloud is this running in" is already a procurement question, not a hypothetical one. Whether Large 4's benchmark numbers hold up against independent evaluation is the normal question any new model release earns -- but the GPU-sourcing claim is the one that's actually novel here, and it's verifiable in a way benchmark scores aren't: it's a claim about where the training happened, not about what the model can do.