Mistral Small 4 Opens a 119B Model That Activates About 60B per Token
Mistral open-sourced Small 4 on 7 September 2026 under Apache 2.0, with a 256k context window.
Direct answer
On 7 September 2026 Mistral released Mistral Small 4 under Apache 2.0. It is a 119 billion parameter mixture-of-experts model that activates about 60 billion parameters per token and offers a 256k context window.
Mistral open-sourced Mistral Small 4 on 7 September 2026 under the Apache 2.0 license. The launch note calls it the first all-in-one model in the Small line, combining jobs that used to sit in separate Mistral models, and says the company is a founding member of NVIDIA's Nemotron alliance.
The architecture in that report is a mixture of experts with 128 experts and 119 billion parameters in total. Four experts activate on each token, which the note puts at roughly 60 billion active parameters. The context window is 256k tokens. The pitch is the usual open-model trade: a large total size, a smaller active slice, and a license that lets a team run the weights.
What the note does not include
The 7 September write-up does not publish an independent benchmark table next to those parameter counts. It describes the model as covering code generation and visual analysis, without scores against a named rival. Until those numbers show up in a primary model card, the concrete specs are the expert count, the 119 billion total, the roughly 60 billion active, the 256k window, and the Apache 2.0 license.
For teams comparing open weights this month, Small 4 is a size-and-license release first. Treat the all-in-one claim as Mistral's positioning until a public eval confirms it.
Source: xix.ai, 7 September 2026.
FAQ
- How many experts fire?
- The report says 128 experts, with 4 active per token.
- What else did Mistral announce that day?
- It said Mistral joined NVIDIA's Nemotron alliance as a founding member.