Mistral has released a preview of Mistral Large 4, a new flagship model available for now only through its API. According to Simon Willison, Mistral says open weights will follow “at the end of this month”. The release was accompanied same-day by version 0.16 of the llm-mistral plugin, which adds support for the model’s reasoning modes.
What was announced
Mistral Large 4 is a mixture-of-experts model with 1 trillion total parameters and 49 billion active parameters per token, according to Simon Willison’s writeup. Mistral trained it on its own cluster of 3,800 NVIDIA Grace Blackwell GPUs, per the same report.
The API preview exposes two reasoning levels, “none” and “high” โ there is no intermediate setting. On the Artificial Analysis benchmark, Willison reports the model scoring 38, just behind DeepSeek 4.1 Flash, a 552 billion parameter model. He contrasts this with Mistral Large 3, released last December, which scored 9 on the same benchmark โ a substantial jump for Mistral between generations.
Separately, llm-mistral 0.16 was published to add support for reasoning models generally, with Mistral Large 4 named as the model that prompted the update. This gives users of Simon Willison’s llm command-line tool a way to call the new model’s reasoning modes through the Mistral API.
- Parameters: 1 trillion total, 49 billion active (MoE)
- Training hardware: 3,800 NVIDIA Grace Blackwell GPUs (Mistral’s own cluster)
- Availability now: API preview only
- Open weights: promised for “end of this month” by Mistral, per Willison
- Reasoning modes: “none” and “high” only
- Benchmark: scores 38 on Artificial Analysis, versus 9 for Mistral Large 3
- Tooling:
llm-mistral0.16 adds reasoning-model support
What it means if you run models locally
For now, nothing changes for anyone running models on their own hardware โ this is an API-only preview, and the sources do not give pricing or rate limits for it. The part that matters to local users is the promised open-weights release. A 1 trillion parameter MoE model with 49 billion active parameters is a large file even before quantisation, and it is likely this will put it well outside what most consumer GPUs can hold at full precision โ this is reasoning, not a figure from Mistral. Once weights are published, quantised versions will probably follow from the usual community sources, and tools like the quantisation explainer and VRAM calculator will be relevant for working out what a given machine can run.
The fact that llm-mistral 0.16 already supports the model’s reasoning modes means anyone using the llm tool with an API key can try Mistral Large 4 today without waiting for local weights, if they are comfortable using a hosted model rather than one running on their own hardware. That is a different use case from fully local, offline inference covered by pages such as local LLMs and the Ollama guide, but it is a quick way to evaluate the model’s quality before deciding whether to wait for self-hosted weights.
What is not known yet
Mistral has not published hardware requirements for running Mistral Large 4 locally, and the sources do not state what quantised sizes, if any, will be offered alongside the open weights. There is no confirmed exact release date beyond “end of this month”, no licence terms for the open-weights version, and no API pricing mentioned in the sources. It is also not stated whether the open release will include the same architecture and active-parameter count as the API preview, or whether context length, supported languages, or multimodal capabilities differ from Mistral Large 3. Anyone planning to run this on their own machine will need to wait for Mistral’s formal weights announcement for those details.
Sources
- Introducing Mistral Large 4: Le chonk — Simon Willison
- Mistral Large 4 — Simon Willison
- llm-mistral 0.16 — Simon Willison
This is a news brief compiled with AI assistance from the sources above. Nothing in it has been tested on my own hardware; when I do test something, I say so and show the numbers.
