Editorial illustration for Microsoft AI's Decision Model Sorts Xbox Feedback 14x Faster Than GPT6
Microsoft's Decision Model Beats GPT-6 Speed by 14x
Microsoft AI's Decision Model Sorts Xbox Feedback 14x Faster Than GPT6
Microsoft put out a new model this week that doesn't write anything at all. Microsoft-Decision-1 takes a situation, a question, and a fixed list of answer options, then hands back a probability score for each one. No paragraphs, no chatty explanations, just numbers a piece of software can act on right away. It's built on top of Alibaba's Qwen3.5-9B, post-trained for this one narrow job, and it's live now in Microsoft Foundry and on OpenRouter through Azure.
The pitch is speed and cost rather than cleverness. Microsoft ran it against 35 other systems across 36 benchmarks and says it came out on top for average accuracy, at 85 milliseconds median latency and $0.042 per million input tokens, with output essentially free. That combination matters for the kind of work this model is aimed at: routing tickets, classifying content, verifying agent outputs, deciding what an automated pipeline should do next.
Microsoft is calling this a new category of model entirely, separate from the text-generation systems most people think of when they hear "AI." Whether that framing holds up depends on how the model performs outside Microsoft's own test results, which is where some of the real questions start.
Microsoft compared 9 systems across 36 benchmarks with 147,137 questions. Benchmarks were kept blind from training. Microsoft-Decision-1 led on average accuracy at 83.5%.
Why this matters
A 9B-parameter model that just scores fixed answers instead of generating text is a narrow tool, but narrow is the point. Most production AI work isn't open-ended chat, it's classification: is this ticket urgent, does this output pass QA, should this agent take step A or B. Microsoft built Microsoft-Decision-1 for exactly that slice, and the Xbox and Copilot numbers (14x faster, 200x cheaper than GPT-6 Sol) suggest there's real money sitting in the gap between "general LLM" and "task-specific scorer." For developers and founders, that's the actionable signal: if your pipeline uses a giant chat model to pick from a short list of options, you're probably overpaying for capability you don't use.
The 46x consistency claim over LLM scoring also hints at something researchers should poke at, calibrated probability outputs may be more trustworthy for verification and routing than sampled text ever was. Worth watching: whether Microsoft opens up weights or benchmarks beyond its own internal tools, and whether rivals ship comparable decision-only models rather than just bigger generalists.
Common Questions Answered
How does Microsoft-Decision-1 differ from general-purpose LLMs like GPT-6?
Microsoft-Decision-1 is a specialized decision-scoring model that doesn't generate text but instead evaluates fixed answer options and returns probability scores for each one. Unlike general LLMs that produce paragraphs and explanations, this model provides immediate numerical outputs that software can act on directly, making it optimized for classification tasks rather than open-ended generation.
What is the performance advantage of Microsoft-Decision-1 for Xbox feedback processing?
Microsoft-Decision-1 processes Xbox feedback 14 times faster than GPT-6 Sol while being 200 times cheaper, according to Microsoft's benchmarks. The model achieved an average accuracy of 83.5% when tested across 36 benchmarks with 147,137 questions, outperforming 8 other systems in the comparison.
What base model is Microsoft-Decision-1 built on and how was it customized?
Microsoft-Decision-1 is built on top of Alibaba's Qwen3.5-9B model and has been post-trained specifically for the narrow task of decision scoring. This specialized post-training allows the 9-billion parameter model to excel at classification and decision-making tasks rather than general-purpose language generation.
What are the primary use cases for Microsoft-Decision-1 in production environments?
Microsoft-Decision-1 is designed for classification tasks that represent most production AI work, such as determining if a support ticket is urgent, evaluating whether output passes quality assurance checks, or deciding which action an agent should take next. The model's speed and cost efficiency make it ideal for these narrow, high-volume decision-making scenarios rather than open-ended conversational tasks.
Where can developers access Microsoft-Decision-1?
Microsoft-Decision-1 is currently available through Microsoft Foundry and on OpenRouter via Azure, making it accessible to developers building classification and decision-scoring systems. This availability allows teams to integrate the specialized model into their production workflows immediately.
Further Reading
- Microsoft-Decision-1: Our model for fast decision-making - Microsoft
- Microsoft's new AI model scores decisions instead of writing text - The Daily Star
- Deploy and use Microsoft-Decision-1 in Microsoft Foundry - Microsoft Learn
- Microsoft puts a 9B decision model in Foundry for agent workflows - RuntimeWire
- Microsoft Launches New AI Model That Makes Decisions 35 Times Faster Than GPT-6 Sol - Business Insider