Mistral AI Releases Mistral Large 4 (Le Chonk): A 1.05T Parameter Multimodal MoE Model
Mistral Large 4, nicknamed Le Chonk, went into public preview today. 1.05T total parameters, 49B active per token, 1M context, native image input. Trained from scratch on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters.
Here is what stood out to me after reading the release and the docs:
→ Only about 4.7% of the weights fire per token
→ That is how a 1T class model serves at $1.36 per 1M input and $4.18 per 1M output tokens
→ The full 1.05T still has to sit in memory, so plan hardware around total, not active
→ 93% on Cybench, 82% on CyberGym-E2E, top 5 on the Artificial Analysis Cyber Index
→ 61.7% DeepSWE v1.1, 28.3% Terminal-Bench 4.0, 49.8% on the Coding Agent Index
→ API is live today as mistral-large-4
I broke down the architecture, benchmarks and a full comparison against DeepSeek V4 Pro, Kimi K3 and GLM-5.3, with an interactive explainer.
Here is the full analysis: https://www.marktechpost.com/2026/10/06/mistral-ai-releases-mistral-large-4-le-chonk-a-1-05t-parameter-open-weight-multimodal-moe/
Technical details: https://mistral.ai/news/mistral-large-4/
Post #1560
489

- 👍 4