Skip to main content
Mistral AI's "Le Chonk" model, a large language model with 1.05 trillion parameters, is released.

Editorial illustration for Mistral AI Releases "Le Chonk": A 1.05 Trillion Parameter Open Model

Mistral AI Releases 1.05T Parameter Open Model

• 4 min read

Mistral AI put a number on the table Monday that's hard to ignore: 1.05 trillion parameters. The French lab released Mistral Large 4, nicknamed "Le Chonk" internally, as a public preview on October 6, 2026. It's a Mixture of Experts model with 49 billion parameters active per token, a 1.6 billion parameter vision encoder bolted on for native image input, and a context window stretching to 1 million tokens.

The training story matters as much as the spec sheet. Mistral built this from scratch on 3,800 Nvidia Grace Blackwell GPUs, run inside the company's own European datacenters rather than rented cloud capacity. That's a statement about infrastructure independence as much as it is about model size.

The API is live now, priced at $1.36 per million input tokens and $4.18 per million output tokens. The open weights, the part that makes this an open-weight release in practice rather than just in name, won't ship until the end of October. Until then, nobody outside Mistral can self-host it, no matter how the pricing looks.

The API is live now at $1.36 per 1M input and $4.18 per 1M output tokens, but the weights do not ship until end of October, so self hosting is not yet possible. Its standout results are in cybersecurity, where Mistral reports 93% on Cybench and 82% on CyberGym-E2E and notes that several closed frontier models score near zero because they refuse the task.

Why this matters

For teams sizing hardware, the 1.05 trillion parameter count is the number that actually bites. The 49 billion active parameters tell you what a single forward pass costs in compute, but someone still has to fit the full model in memory before any of that routing math happens. That's a real constraint for anyone outside a hyperscaler's data center, and Mistral hasn't said how ML4 is meant to be served at smaller scale.

We'd also note the missing details: no expert count, no routing scheme, no layer layout. Those aren't trivia. They determine whether this is genuinely efficient or just a trillion-parameter model wearing a mid-tier price tag.

Mistral training on 3,800 Grace Blackwell GPUs in its own European centers signals real infrastructure investment and a push for sovereignty from US cloud providers, which matters for European developers and regulators watching that story. Until the architecture specifics land, though, treat the "efficient trillion-parameter model" framing as a claim to verify, not a settled fact. Watch for the full technical report.

Common Questions Answered

What are the key technical specifications of Mistral Large 4 'Le Chonk'?

Mistral Large 4 is a 1.05 trillion parameter Mixture of Experts model with 49 billion parameters active per token, featuring a 1.6 billion parameter vision encoder for native image input and a context window of 1 million tokens. This architecture allows the model to handle both text and image inputs while maintaining efficient compute during inference through its mixture of experts routing mechanism.

When will the model weights for Le Chonk be available for self-hosting?

The model weights for Mistral Large 4 are scheduled to ship at the end of October 2026, though the API is already live as of October 6, 2026. Currently, self-hosting is not yet possible, but users can access the model through Mistral's API at $1.36 per 1M input tokens and $4.18 per 1M output tokens.

What are Le Chonk's standout performance results compared to other models?

Mistral Large 4 achieves exceptional results in cybersecurity tasks, scoring 93% on Cybench and 82% on CyberGym-E2E benchmarks. Notably, several closed frontier models score near zero on these same cybersecurity tasks because they refuse to complete them, making Le Chonk's performance particularly impressive in this specialized domain.

What hardware challenges does the 1.05 trillion parameter size present for deployment?

The full 1.05 trillion parameter model must fit entirely in memory before any routing operations can occur, which represents a significant constraint for teams outside hyperscaler data centers. While the 49 billion active parameters per token indicate the compute cost of a single forward pass, the full model size requires substantial hardware resources that Mistral has not yet detailed for smaller-scale deployment scenarios.

LIVE21:08Artificial' Warns of AI Future Controlled by Tech Giants or China