Xiaomi opens MiMo-V2.6 under MIT license — and tops the ranking for open-weight models
Xiaomi has released the MiMo-V2.6 series as MIT-licensed open-weight models and claims the top spot among open models on the Artificial Analysis Intelligence Index — 46 points, ahead of GLM-5.3 at 45 and Kimi K3 at 44. At the same time, the company publishes the training costs of the reinforcement learning run, over 7,000 RL environments, and the framework itself — an unusual degree of openness in an industry where such figures are rarely made public.
The news
Xiaomi launched and opened the MiMo-V2.6 series on September 22, 2026 (Artificial Analysis lists September 21 as the release date). The series consists of the multimodal flagship model MiMo-V2.6-Pro, the faster MiMo-V2.6-Flash, and the distilled MiMo-V2.6-Distill-Qwen-9B, along with research resources for reinforcement learning (RL) (TechNode).
What makes the release remarkable is not just the capability level but the degree of documentation: Xiaomi has published the training costs of the RL run, reportedly streamed the training dashboard live, and opened both the technical report and the environments and code. The company itself highlights the openness as what sets the release apart from most contemporary model drops — enough published material for the claims to be tested by others.
The models
According to the model card from Xiaomi's MiMo team, MiMo-V2.6-Pro-RL is a sparse mixture-of-experts model with 1.02 trillion total parameters, of which 42 billion are activated per token. It supports 1 million tokens of context and takes text, image, video, and audio as input with text as output, under an MIT license (Unite.AI).
The pricing is aggressive: Xiaomi charges $0.435 per million uncached input tokens and $0.87 per million output tokens for Pro, while Flash costs $0.14/$0.28. Artificial Analysis measures the model at $0.13 per index task and around 134 tokens per second in output speed, and confirms the 1 million token context window with multimedia input (VentureBeat).
The RL run
Xiaomi claims that both models completed 30 RL steps in under six days, with around 750,000 trajectories. The reported costs are $2.62 million for Pro and $850,000 for Flash — a combined total of roughly $3.47 million. The total matches the sum of TechNode's two cost figures, and a Japanese blog also reproduces Xiaomi's own summary of around $3.47 million (TechNode; note.com). All cost figures originate from Xiaomi itself.
The company describes scaling RL compute along three axes (Unite.AI):
- Large asynchronous batches — a fully asynchronous architecture with 1,568 samples per update, trained at up to 1 million context, with 3.5–3.7 billion tokens per step.
- Multi-task suite — coding, general agents, visual, and cybersecurity tasks mixed across multiple frameworks.
- More grader compute — relative comparison within each batch for more precise reward signals.
According to Xiaomi, the average pass rate on the training tasks rose by 25 percent (Flash) and 12 percent (Pro) in relative terms, and the scores on DeepSWE v1.1 — a hold-out benchmark for long-horizon software development — rose from 48.8 to 65.68 for Flash and from 58.4 to 72.57 for Pro. These figures are Xiaomi-reported, not independently verified. There is, incidentally, a small discrepancy in Xiaomi's own numbers: the company's benchmark tables list Pro at 71.9 on DeepSWE, while the RL run's final figure is given as 72.57 — a gap the sources do not explain.
The verification picture
There is one solid independent confirmation: Artificial Analysis' own model page records MiMo-V2.6-Pro at 46 points on the Intelligence Index v4.3.2 — a composite of ten evaluations, including AA-Briefcase v1.1, GDPval-AA v2.1, Terminal-Bench 4.0, SciCode, and Humanity's Last Exam — and places it first among open-weight models, ahead of GLM-5.3 (max) at 45 and Kimi K3 (max) at 44 (Unite.AI). VentureBeat adds that the model is tied with Grok 4.7, which debuted the same day and also scores 46, and that it sits above DeepSeek V4.1 Flash (39) and DeepSeek V4.1 Pro (36) (VentureBeat).
Most other benchmark figures are vendor-reported: Terminal Bench 2.1 at 89.9 percent, CyberGym at 94.0 percent, and AutomationBench v1.0.6 at 53.1 percent are Xiaomi's own results, not independently benchmarked (Forkast via Yahoo Tech).
There is also an unresolved contradiction in the comparison with closed models: TechNode, citing Xiaomi's own table, reports that leading closed models score 53 on the index, while VentureBeat and Unite.AI place Grok 4.7 at the same level (46) as MiMo-V2.6-Pro. The difference likely reflects different index versions or model configurations, but is not resolved in the available source material. Regardless: the flagship model still sits noticeably behind the leading closed systems by Xiaomi's own figures.
What is open
Beyond the MIT-licensed weights, Xiaomi open-sources (Unite.AI; Android Authority):
- A full technical report
- Over 7,000 reinforcement learning environments
- An end-to-end RL framework
- Desktop apps for Windows and macOS
Open questions
Several matters remain uncertain. The margin over GLM-5.3 is one point on one index — a narrow lead, not a clear shift in the hierarchy of open models. The closed-model comparison is inconsistent between sources. And although the published training costs, open-sourced code, and environments make the claims more testable than usual, it remains to be seen whether third parties actually reproduce both the results and the cost figures with the released material. Until then, the release is above all a successful documentation project — and a concrete, though company-reported, data point in the debate over reinforcement learning as the next scaling lever.

