Xiaomi's MiMo-V2.6 Is Now the Strongest Open-Weight Model on the Board
Xiaomi's MiMo-V2.6 Is Now the Strongest Open-Weight Model on the Board
The smartphone company best known for phones and electric cars just dropped the strongest downloadable AI model in the world — and it wants you to run it.
Xiaomi released MiMo-V2.6 on 21 September 2026, with weights on Hugging Face under a permissive MIT license. The family ships as three models: MiMo-V2.6-Pro, its smaller sibling MiMo-V2.6-Flash, and a ~9B distilled variant. All three are free to download, fine-tune, and put in production.
Why the 46.32 score matters
The number Xiaomi is leading with is the Artificial Analysis Intelligence Index, a composite reasoning benchmark. Pro landed at 46.32 at launch — the highest score ever recorded for an open-weight model, edging past Moonshot's Kimi K3 and Alibaba's Qwen3.8 Max to take the crown. That puts a downloadable model on the same board as closed, proprietary frontier systems like GPT-5.6 Sol and Claude Opus 5, with Xiaomi claiming parity with both across most agent benchmarks.
Here is what that means in practice: open weights now sit roughly at the frontier, not a step behind it. For anyone who has been watching intraday model rankings, that is the headline takeaway.
The architecture worth knowing
Pro is a sparse mixture-of-experts model: roughly 1.02 trillion total parameters, but only about 42B activated per token. It routes through 384 experts, with eight firing on any given token. Because sparse MoE switches on only part of the network for each input, running cost tracks the active params rather than the full 1T.
A few details worth flagging: - MiMo-V2.6-Flash holds ~309B total / ~15B activated, with a 1-million-token context window. - The backbone mixes Global Attention and Sliding Window Attention in a 1:5 ratio — full global layers only where quadratic attention earns its keep, sliding window elsewhere to cut memory. - Speculative multi-token prediction (MTP) layers speed up decoding by roughly 2.5-3.7x. - The model was trained with scaled reinforcement learning across 7,000+ tasks, from software engineering to vulnerability reproduction to web work.
API pricing is unchanged from the V2.5 series, which quietly makes this one of the cheapest frontier-class models you can call by API.
What it means for self-hosting
The two flagship models are large — a 1.02T sparse MoE won't fit on a single consumer GPU. The honest practical takeaway for self-hosters is the 9B distilled model, which is dense and light enough to run locally on a single 24GB card. Flash sits in between, viable for labs with modest infrastructure. The MIT license means no gatekeeping on commercial hosting, so expect the usual wave of third-party providers to pick these up fast.
The bigger picture
The significance here is not just the score. Xiaomi is a consumer hardware giant with enormous manufacturing and distribution reach, and it is now competing seriously in frontier open-weight AI against the usual suspects. Open-weight competition keeps getting cheaper and stronger, and the release — weights, training environments, and RL code shipped together — is engineered to be reproduced. For the ecosystem, that matters more than any single leaderboard number. Read more on the technical details at https://tama.fdhcl.com in our sister coverage.
