Editorial illustration for Xiaomi Credits AI Gains to Expanded Reinforcement Learning
Xiaomi's MiMo-V2.6 Tops Open Model Rankings
Xiaomi Credits AI Gains to Expanded Reinforcement Learning
Xiaomi put out its MiMo-V2.6 lineup this week, and the flagship version now sits at the top of the open-model rankings compiled by Artificial Analysis. MiMo-V2.6-Pro scores 46 points on the firm's Intelligence Index, edging out Kimi K3 and Qwen for the best publicly available result. What separates Xiaomi's model from the pack isn't just the score, it's the price tag attached to it.
Input tokens run $0.435 per million, output tokens $0.87 per million, and Artificial Analysis pegs the cost of a single benchmark task at roughly $0.13, far below what comparable models charge. That combination of top-tier performance and low cost places MiMo-V2.6-Pro on the Pareto frontier, the sweet spot where intelligence and price both work in a buyer's favor. Under the hood, Pro is a mixture-of-experts system built on 1.02 trillion parameters, though only 42 billion activate for any given request.
Xiaomi paired it with a leaner sibling, MiMo-V2.6-Flash, for cheaper workloads. Xiaomi says the jump in capability traces back to a much larger reinforcement learning push than it used previously, one the company opened up for outside inspection.
Xiaomi has released its MiMo-V2.6 lineup, and the flagship model tops the charts among openly available models while costing a fraction of the competition per task. The gains come from a massively expanded round of reinforcement learning, though the company also stands accused of borrowing from Anthropic's Claude.
Why this matters
For developers weighing which open model to build on, MiMo-V2.6-Pro's price-to-performance ratio is the headline, and Xiaomi's decision to publish its RL tools and task environments makes that gain reproducible rather than just a benchmark claim. That's the part worth testing yourself before taking Xiaomi's rankings at face value. But the Anthropic allegation complicates the story Xiaomi wants told.
If a leading open model's reinforcement learning gains were partly built on unauthorized use of Claude's outputs, the "we scaled RL three ways and it worked" narrative starts looking less like a training breakthrough and more like a shortcut with a legal cloud over it. Researchers should watch whether Xiaomi's released materials hold up to independent scrutiny, and whether Anthropic pursues this beyond a public accusation. For founders building on open models, the practical question isn't just "is it cheap and capable," it's "will the provenance of its training data become someone else's problem later." Cost and performance are easy to measure.
Clean data lineage, right now, is not.
Common Questions Answered
How does MiMo-V2.6-Pro's pricing compare to competitors in the open-model rankings?
MiMo-V2.6-Pro costs $0.435 per million input tokens and $0.87 per million output tokens, making it significantly cheaper than competing models while maintaining the top position on Artificial Analysis's Intelligence Index. This price-to-performance ratio is a major competitive advantage for developers choosing which open model to build on.
What reinforcement learning approach did Xiaomi use to improve MiMo-V2.6?
Xiaomi credits MiMo-V2.6's performance gains to a massively expanded round of reinforcement learning, which the company has made reproducible by publishing its RL tools and task environments. This approach allows developers to test and verify Xiaomi's improvements rather than relying solely on benchmark claims.
What is the controversy surrounding Xiaomi's MiMo-V2.6 development?
Anthropic has accused Xiaomi of borrowing from Claude during the development of MiMo-V2.6, which complicates the narrative Xiaomi wants to present about its reinforcement learning gains. This allegation raises questions about the independence and originality of the model's development process.
Where does MiMo-V2.6-Pro rank among publicly available AI models?
MiMo-V2.6-Pro scores 46 points on Artificial Analysis's Intelligence Index and sits at the top of the open-model rankings, edging out competitors like Kimi K3 and Qwen. This makes it the best publicly available result according to the firm's evaluation metrics.
Further Reading
- Xiaomi opensource MiMo-V2.6 and the RL machinery behind it - RuntimeWire
- Xiaomi livestreams MiMo-V2.6 reinforcement-learning runs - TechNode
- MiMo-V2.6: Scaling Up Reinforcement Learning for Self-Improvement - Xiaomi MiMo
- Xiaomi Publicly Reveals the RL Training Process of MiMo-V2.6 - Aibase
- Xiaomi's MiMo-V2.6-Pro debuts as the top open-weights model in the world - VentureBeat