Skip to main content
Presenter in a sleek Xiaomi hall gestures to a large screen showing MiMo-V2-Flash AI speed and price, with brand logos behind.

Editorial illustration for Xiaomi Unveils MiMo-V2-Flash AI Model with 150 Tokens/Sec at Low Cost

Xiaomi's MiMo-V2-Flash: 150 Tokens/Sec AI Model Breakthrough

Updated: 3 min read

Xiaomi is selling an AI model for less than the price of a text message. For a tenth of a cent per million input tokens, you can rent MiMo-V2-Flash. It is fast and it is cheap, which is exactly the point.

The numbers are a direct challenge. It runs at 150 tokens per second. It matches or beats established heavyweights like DeepSeek V3.2 and Moonshot's Kimi K2 on standard reasoning tests.

On long-context work, it beats Kimi. Its most telling score is 73.4% on SWE-Bench Verified, a benchmark for practical coding tasks. That figure leaves every other open-source model behind and sits just behind OpenAI's proprietary GPT-5-High.

Xiaomi claims it rivals Anthropic's Claude 4.5 Sonnet on coding for a tiny fraction of the cost.

According to Xiaomi, the model delivers inference speeds of up to 150 tokens per second and operates at a low cost of $0.1 per million input tokens and $0.3 per million output tokens. On benchmarks, Xiaomi claimed MiMo-V2-Flash achieves performance comparable to Moonshot AI's Kimi K2 Thinking and DeepSeek V3.2 across most reasoning tests, while surpassing Kimi K2 in long-context evaluations. In agentic tasks, the model scored 73.4% on SWE-Bench Verified, outperforming all open-source peers and approaching OpenAI's GPT-5-High.

Xiaomi also said it matches Anthropic's Claude 4.5 Sonnet on coding tasks at a fraction of the cost. MiMo-V2-Flash uses a Mixture-of-Experts architecture to split large neural networks and has 309 billion parameters, allowing it to balance performance and efficiency. It allows Xiaomi engineers to work on architectural optimisations, significantly reducing the cost of processing long prompts by limiting how much past context the model needs to re-evaluate.

Luo Fuli, a former DeepSeek researcher who recently joined Xiaomi's MiMo team, described the release as "step two on our AGI roadmap" in a post on X, referring to artificial general intelligence.

This is not a side project. It is a 309-billion-parameter Mixture-of-Experts model designed to make long-context processing affordable. The architecture limits the computational cost of revisiting past context, turning a major expense into something manageable.

The subtext matters. Luo Fuli, a key researcher who recently left DeepSeek for Xiaomi, framed this release as step two on an AGI roadmap. The company known for phones and appliances is now setting the price floor for serious AI inference. They are using open-source code to apply pressure where it hurts most: the bill.

Common Questions Answered

How fast is the inference speed of Xiaomi's MiMo-V2-Flash AI model?

The MiMo-V2-Flash AI model delivers an impressive inference speed of up to 150 tokens per second. This high-performance capability positions the model as a competitive option in the rapidly evolving open-source AI landscape.

What are the pricing details for Xiaomi's MiMo-V2-Flash model?

Xiaomi offers the MiMo-V2-Flash at a remarkably low cost of $0.1 per million input tokens and $0.3 per million output tokens. These competitive pricing rates are designed to make advanced AI technology more accessible to developers and businesses.

How does MiMo-V2-Flash perform on benchmarks compared to other AI models?

According to Xiaomi, the MiMo-V2-Flash achieves performance comparable to Moonshot AI's Kimi K2 Thinking and DeepSeek V3.2 across most reasoning tests. The model particularly excels in long-context evaluations and scored an impressive 73.4% on SWE-Bench Verified in agentic tasks, outperforming other open-source peers.

LIVE22:59Sam Altman Addresses AI Alarm Over Autonomous Agents