Microsoft on Friday launched Microsoft-Decision-1, a specialized model for fast, structured decision-making, on its Foundry platform — and built it on Alibaba's open-weight Qwen3.5-9B rather than on any model from its close partner OpenAI, The New Stack reported.
The launch lands three days after OpenAI opened its Decisions API to every developer in public beta, and on the same day TypeSafe, the startup that ignited the decision-model category with Jev roughly three and a half weeks ago, announced an $870 million Series A at a $7.5 billion valuation led by a16z. Microsoft priced Decision-1 at exactly Jev's rate: $0.042 per million input tokens with free output. For more on the accelerating race to make AI cheaper and faster, see our ongoing AI industry coverage.
Priced to Match Jev, Built on a Rival's Open Weights
Decision-1 is a post-trained model: Microsoft started from Alibaba's Qwen3.5-9B and tuned it for structured decision tasks rather than open-ended generation. The company says it plans to rebase the model on its own MAI models as well as OpenAI's, but the initial choice of a Chinese open-weight base is notable for a company that invests billions in OpenAI.
Microsoft is also far from the first mover. In the three and a half weeks since Jev launched, OpenAI, Upstage, Perplexity, Cloudflare, and AWS have all shipped decision models of their own, according to The New Stack. The Qwen3.5 base Microsoft picked has become something of an industry default: Cognition's vice president of engineering, Jared Palmer, spent about $95 in Modal H100 compute porting his open-source Kev models to Qwen3.5, and Cloudflare built its Clef-flash model on the same Qwen3.5-9B base.
Price is converging too. Palmer lists Kev-4B on OpenRouter at $0.042 per million input tokens, Perplexity charges $0.02, and OpenAI charges more than double Jev's rate. At those prices, revenue from individual decision calls is minimal — which suggests Microsoft's bigger opportunity is keeping agent traffic, including the generative calls around each decision, running through Foundry.
Microsoft Is Its Own First Customer
Microsoft chairman and CEO Satya Nadella announced the model on X on Friday, writing that Decision-1 "delivers top performance on structured decision tasks, outperforming both LLMs and other decision models in latency and quality," and adding, "We're already testing it across Microsoft."
Four internal teams backed that up, per The New Stack. Xbox Research used Decision-1 to sort more than 10,000 pieces of player feedback. The Copilot team graded chat and agent responses with it. On-call engineers used it to pull context during live incidents, and Microsoft Discovery used it to score experiments before an agent replans.
Microsoft's internal numbers, as reported, claim the model is more than 14 times as fast as GPT-6 Sol at a fraction of the cost for the Xbox workload, and 46 times more consistent in Discovery. Achint Srivastava, vice president of software engineering in Microsoft's Office of the CTO, introduced Decision-1 as a way to add decision-making to existing applications, agents, and workflows "in a secure, trusted environment."
What Decision-1 Cannot Do
The launch also shows where Microsoft trails rivals. Decision-1's Foundry listing says it accepts up to 32,768 tokens of text and returns JSON, but it does not support images — unlike OpenAI's Decisions API and Cloudflare's Clef, which uses a vision encoder. And while Cloudflare released Clef under the permissive Apache 2.0 license, Microsoft has not announced open weights for Decision-1.
There is also a compatibility question. AWS, Upstage, and Ollama have adopted TypeSafe's System One API, now a common interface across the category, but Microsoft has not said whether Decision-1 is fully compatible — though its Foundry sample code calls a /systemone endpoint.
Calibration claims deserve scrutiny as well. Microsoft says Decision-1's probabilities are calibrated, meaning a 90% prediction should be right about nine times in ten on representative cases, and the company's own documentation advises customers to validate calibration on their data. That caution echoes the JevOut preprint, in which USC computer science researcher Zixiang Xu and co-authors found that short, natural-sounding additions to the context flipped 312 of Jev's 508 initially correct decisions — and in 229 cases Jev assigned at least 70% probability to the wrong answer. Three other scoring systems showed flip rates between 64.9% and 73.2% under the same testing approach. Microsoft says it tested Decision-1 with eight kinds of perturbations, including reordered options and paraphrased descriptions.
Routing Decisions Across Copilot, GitHub, and Xbox
The model may have a bigger job ahead. On Wednesday, Microsoft said GitHub Copilot will soon decide when to run a task on-device and when to send it to cloud-scale models, though it has not disclosed what Copilot sends to the cloud. Model routing is among the use cases Microsoft lists for Decision-1, but the company has not confirmed that the model will make those calls. With routing decisions to make across Copilot, GitHub, and Xbox, Microsoft has an incentive to handle them in-house.
The competitive stakes are visible in TypeSafe's traction. TypeSafe CEO Diogo Almeida posted on X that "29.4% of the Fortune 500 showed up" in the three weeks since Jev launched, that the model "accidentally served trillions of tokens a day," and, with an eye on his team's workload, that "it's been 3 weeks we would like to sleep now." Those are the same enterprises Microsoft sells Azure to — which makes Decision-1's arrival less a bet on a new market than a defense of an existing one.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →

