Mistral Promises Open-Weight Release of ML4 on October 27 – Benchmarks Still Unverified

Mistral AI unveiled ML4, nicknamed "Le Chonk," in public preview on Tuesday – and promised to release the model weights on October 27. The launch came one day after the American company Reflection AI announced its open Beam model, and just…

A large heavy metal sphere rests on a concrete floor in an empty industrial hall beside an open unlocked cage door – an illustration of a heavy AI model whose weights are soon to be released openly.
Illustration
Gift article

Mistral Promises Open-Weight Release of ML4 on October 27 – Benchmarks Still Unverified

Mistral AI unveiled ML4, nicknamed "Le Chonk," in public preview on Tuesday – and promised to release the model weights on October 27. The launch came one day after the American company Reflection AI announced its open Beam model, and just weeks after Mistral raised €3 billion at a valuation above €21 billion. But while the company claims ML4 will be the most capable open-weight model outside China, all performance claims so far are Mistral's own – the final benchmark scores are still missing.

What Happened

The French AI champion Mistral AI launched a public preview on Tuesday of a new frontier model that the company itself says will be the most capable open-weight model developed outside China when the weights are released (Euronews via MSN, source trace bacf04ea).

The model has the nickname "Le Chonk" – "the big, chunky one." Chief scientist and co-founder Guillaume Lample called it "a new generation of models" at a press conference attended by Euronews (bacf04ea).

ML4 is available now through Mistral's API on Mistral Studio. The weights are planned for release on October 27 (bacf04ea, and ZDNET, source trace 2f5a0f26).

The timing is no accident. The day before – October 5, according to TechCrunch's dateline – Reflection AI debuted Beam, an open-weight model aimed explicitly at Chinese models, and at Western players including Mistral (TechCrunch, source trace e2c1d236). The open-weight race has thus gained two new Western challengers in two days.

What Mistral Claims – and What Is Actually Verified

Almost all performance statements about ML4 come from Mistral itself and should be read as marketing until independent verification exists. Here are the company's claims as reported:

  • Trained from scratch: According to Mistral, ML4 was trained from zero over two months on 4,000 NVIDIA Grace Blackwell GPUs in the company's own European data centers (bacf04ea).
  • Languages: The model is said to cover more than 160 languages, including all official EU languages, and to work in both Latin and non-Latin scripts (bacf04ea). This is a direct appeal to European public agencies, which often operate in languages the big American models have historically covered unevenly.
  • Performance: Mistral says ML4 leads among open-weight models on production and finance tasks as well as multimodal use, that in coding it closes the gap to other frontier models, and that it has the best results of any model – including closed ones – on grounding properties, plus top scores on semiconductor benchmarks (bacf04ea). None of these claims are independently verified.
  • Cybersecurity: A Mistral spokesperson wrote in an email to ZDNET that ML4 surpasses the best models from Kimi, DeepSeek and Meta "in absolute terms on cyber skills" (2f5a0f26).
  • Positioning against China: Pierre Stock, vice president of science at Mistral, said that all AI companies are moving "at full speed," that the three-year-old company "is definitely closing the gap" to competitors, and that the new model is in some respects "stronger than China's models from the summer" (bacf04ea).

The most important caveat: Mistral itself has said it is waiting on final benchmark scores (2f5a0f26). ZDNET references an early third-party analysis that purportedly shows ML4 on par with more expensive proprietary models, but the analysis is unnamed and cannot be weighed as independent confirmation. Until October 27, and until third-party benchmarks exist, "best outside China" stands as a promise – not a fact.

The Cybersecurity Angle: Ownership as an Argument

The most concrete positioning of ML4 concerns cybersecurity. In a written statement to ZDNET, Mistral wrote:

"ML4 is the beginning of a leading generation of open, customizable cybersecurity models that enterprises can fully own and control, with no vendor lock-in," adding: "Enterprises and states should not have to rely on a closed model vendor that can arbitrarily switch off their cyber defense capabilities" (2f5a0f26).

Lample tied this to the threat landscape: "Cyber defense capabilities will enable enterprises and governments to defend themselves against threat actors who jailbreak closed models to carry out cyberattacks," he told ZDNET (2f5a0f26).

In practice, this means something different from ordinary API access. Early test partners get access to a version of ML4 with fewer safety guardrails and "extended cyber skills" – an approach ZDNET compares to Anthropic's Project Glasswing and OpenAI's rollout of Astra (2f5a0f26; the comparison is ZDNET's, not Mistral's own statement). For a company or a government ministry that needs cyber defense with data and weights under its own control, this is a fundamental difference from consuming a closed API.

The argument also has a flip side. Open models are, as CBS News has covered the debate, more vulnerable to jailbreaking and harder to oversee than closed models (source trace e9bfaef7). But the counterargument was laid out in a joint letter from July of this year, signed by more than 70 companies including Google, Microsoft and NVIDIA: open models "expand defense capabilities, increase transparency, and enable vulnerabilities to be discovered and patched" (e9bfaef7). ML4's cybersecurity positioning places Mistral squarely in the second camp of this question – open weights as defense, not merely attack surface.

The Context: Capital, Beam, and the Open-Weight Race

The ML4 release comes hard on the heels of two other events that frame it.

First, the funding. In September, Mistral announced it had raised €3 billion at a valuation above €21 billion – according to the company itself, the largest equity round ever completed by a European technology company (bacf04ea). That provides the capital base both for training in its own data centers and for the broader push.

Second, the competition. Reflection's Beam is a mixture-of-experts model with 501 billion parameters and 23 billion active parameters, pre-trained on 23.8 trillion tokens, with a one-million-token context window (e2c1d236). Reflection says Beam scores on par with Z.ai's GLM-5.2 on advanced reasoning benchmarks and surpasses today's leading Western open models, while using "3–4 times less inference compute" – claims TechCrunch explicitly flags as independently unverified (e2c1d236). The company, founded in 2024 by two former Google DeepMind researchers, has raised roughly $4.7 billion from investors including Nvidia, Sequoia Capital and Lightspeed Venture Partners, according to PitchBook (e2c1d236).

The comparison reveals a strategic divide: Reflection competes primarily on compute efficiency and positions itself against both Chinese and Western open models, while Mistral combines performance claims with an explicit European sovereignty narrative – its own data centers, more than 160 languages, and a cybersecurity argument aimed at states and critical infrastructure. Beam and ML4 thus attack the same problem space, but from different angles.

What October 27 Must Show

The weight release on October 27 is where the promises come due, and at least four questions will determine whether they are met:

The benchmark scores. Mistral itself has said final benchmarks are still missing (2f5a0f26). Until third parties verify the claims on grounding, semiconductors, production, finance and cyber skills, they remain exactly that – claims.

The model's actual size. The sources give no consistent answer on ML4's parameter count, and it cannot be settled from the available evidence. For an open-weight release this is no detail: size determines who can actually run the model locally, and therefore who "full ownership" is really for.

The jailbreak risk in practice. A less-guardrailed version for test partners follows a known pattern from Anthropic and OpenAI, but for a model marketed on cyber skills to enterprises and states, the question is whether the safety architecture can withstand the weights becoming publicly available – which is exactly what happens on October 27.

Does "European sovereignty" mean anything operationally? Training in European data centers and coverage of all EU languages are tangible facts, insofar as they are confirmed. But the sovereignty story also rests on the weights actually being released, the license permitting genuine customization, and the model being good enough that European organizations choose it over closed American alternatives.

What Stock declared at the press conference – that the company "is definitely closing the gap" (bacf04ea) – is precisely worded. The gap is still being closed. The proof arrives on October 27.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.