Qatar AI Institute
← Back to articles

7 October 2026 · By

Mistral Large 4: A Trillion-Parameter Open-Weight Model

Mistral Large 4: A Trillion-Parameter Open-Weight Model

On 6 October 2026 the French AI company Mistral introduced Mistral Large 4, a model with about one trillion parameters in total, of which only 49 billion are used for any single token. A public preview is available now, and Mistral says the full open weights will be released by the end of October 2026. It is a useful moment to understand a design called mixture of experts, which is how very large models stay affordable to run.

What Mistral announced

  • Size: about 1 trillion total parameters, 49 billion active per token.
  • Type: a "natively multimodal" model that takes text and images, and combines instruction-following and reasoning in one model.
  • Training: on 3,800 NVIDIA Grace Blackwell GPUs in Mistral's own European datacenters, on multilingual data covering more than 160 languages, including all official EU languages.
  • Availability: a public preview through Mistral Studio now; open weights, designed for self-deployment, promised by the end of October.
  • Caveat from Mistral: this is a preview that is still being refined. Mistral says it has seen no signs of saturation, which suggests the final version may improve.

For context, Mistral's previous flagship, Mistral Large 3 (December 2025), was reported as a mixture-of-experts model with 675 billion total and 41 billion active parameters. Large 4 is larger in total, and more of it is active per token.

What "mixture of experts" means

In an ordinary ("dense") Transformer (see What Is a Transformer?), every parameter takes part in processing every token. In a mixture-of-experts (MoE) model, the large feed-forward blocks are split into many smaller "experts". A small network called the router looks at each token and sends it to just a few of them. The others do nothing for that token.

Diagram: a token enters a router, which sends it to three highlighted experts out of sixteen. The other experts stay idle.
A router activates only a few experts per token. Simplified and illustrative; real models have far more experts.

The result is a model with the knowledge capacity of a trillion parameters but a per-token compute cost closer to that of a 49-billion-parameter model, roughly 5% of the total. This is why "total parameters" and "active parameters" are listed separately, and why comparing a 1T MoE model with a 70B dense model by parameter count alone is misleading. The catch is memory: to serve the model you still have to hold all the experts, so the hardware requirements follow the total size.

Why it matters

  • Open weights at frontier scale. Many of the largest models are available only through an API. If Mistral releases the weights as promised, universities and companies could run and study a trillion-parameter model themselves, including on their own infrastructure, which matters for data-sensitive work in the Middle East and elsewhere.
  • European compute. The model was trained in European datacenters, part of a broader push to build AI capacity outside the United States and China.
  • Multilingual focus. Coverage of more than 160 languages is relevant for regions where Arabic and other languages are under-served by English-centred models. How well it handles Arabic is something to test once it is available.

What to be careful about

Most of what is public so far comes from Mistral's own announcement. Independent benchmark results, how the preview compares with the final weights, and the exact licence terms will only be clear after release, so treat claims about quality as provisional until they have been tested by others.

What it means if you are learning AI

MoE is now a standard part of frontier model design, so it is worth understanding routing, load balancing between experts and the memory-versus-compute trade-off. The best starting point is a solid grounding in Transformers and deep learning: see Neural Network Architectures: Transformers, GNNs and GANs, and for another recent idea about spending computation unevenly, read our explainer on ALoDLM. Our Introduction to Artificial Intelligence course is a good place to build the foundations.

What to watch next

  • Whether the open weights arrive by the end of October, and under what licence.
  • Independent benchmarks, including on Arabic.
  • How much hardware is needed to self-host a model of this size.

Source: Mistral AI, "Introducing Mistral Large 4", 6 October 2026. Background on Mistral Large 3 from press coverage of its December 2025 release. This article is an independent explainer based on the company's announcement.

Related articles