Microsoft is quietly replacing AI models from OpenAI and Anthropic with its own in-house models inside Copilot, routing tens of thousands of queries each week through homegrown software in products such as Excel and Outlook. The shift, reported on July 7, 2026 by PYMNTS and The Decoder, signals a concerted effort by the company to cut the cost of running its flagship AI assistant — even at the risk of delivering slightly less capable results to some users.
The move is striking given Microsoft's deep ties to OpenAI, in which it has invested billions. But under AI chief Mustafa Suleyman, Microsoft has been steadily building its own model family — branded MAI — and now appears determined to use it to reduce reliance on the very partners that helped make Copilot a household name. For more context on this story, see our ongoing more AI stories.
Tens of thousands of queries rerouted
According to The Decoder, Microsoft is already running tens of thousands of Copilot queries per week through its own MAI models, particularly for routine tasks embedded in Office applications like Excel and Outlook. PYMNTS reported that the company has shifted thousands of prompts in its Office products to the in-house models, describing the change as a quiet but deliberate campaign to bring more of Copilot's workload in-house.
The rationale is primarily financial. Every query sent to a third-party model carries a per-token cost, and with hundreds of millions of potential Copilot users across Microsoft's productivity suite, those charges scale rapidly. By substituting its own models for high-volume, lower-stakes tasks, Microsoft can keep the most expensive third-party calls for the queries where they matter most.
Suleyman's mandate to 'eliminate' external costs
The strategy carries the imprint of Mustafa Suleyman, Microsoft's chief of consumer AI, who The Decoder reported wants to "ultimately eliminate" the cost of external models. That ambition reflects a broader industry reality: as AI assistants move from novelty to everyday utility, the economics of inference — the computing required to generate responses — have become a decisive battleground.
For Copilot customers, the trade-off could be tangible. The Decoder noted that relying more heavily on in-house models may mean "less performance for the same price," since Microsoft's MAI family, while improving, may not yet match the top-tier capabilities of the most advanced OpenAI and Anthropic models on every task. Microsoft appears to be betting that for routine productivity work — summarizing an email, formatting a spreadsheet, drafting a reply — a cheaper in-house model is "good enough."
Building the MAI model family
The shift did not happen overnight. Microsoft has spent years assembling its own model capabilities under the MAI banner, investing in research talent and the massive data-center infrastructure required to train and serve models at scale. That buildup now gives the company something it lacked in Copilot's early days: a credible internal alternative to the frontier models it buys from partners.
The appeal of running on homegrown models extends well beyond per-query savings. When Microsoft relies on its own software, it controls the update cadence, the safety tuning, and the data handling — an increasingly important consideration for enterprise customers worried about where their information flows. It also removes a layer of vendor dependency that has grown more conspicuous as the cost of external models has climbed.
A recalibrated relationship with OpenAI
The shift also reframes Microsoft's relationship with OpenAI, long the public face of Copilot's intelligence. Microsoft remains deeply intertwined with OpenAI through investment, cloud infrastructure, and product integration, and the companies continue to collaborate on frontier research. But the move to route everyday Copilot traffic through MAI models shows Microsoft hedging its dependence, treating OpenAI less as an indispensable engine and more as one supplier in a growing portfolio.
The same logic applies to Anthropic, whose Claude models Microsoft has also offered through its platforms. As per-token pricing from leading US providers has climbed, several companies have begun exploring cheaper alternatives, including increasingly capable open-source and Chinese models — a trend CNBC has documented as Chinese AI systems gain ground on cost.
The inference-cost wars
Microsoft's gambit is an early salvo in what is shaping up to be a defining contest of the AI era: the war over inference costs. As analysts have warned, the expense of running AI at scale can rival or exceed the cost of the engineers building the software, pressuring every company that embeds AI into mass-market products to find efficiencies.
For Microsoft, owning more of the model stack — and the margins that come with it — is the natural endpoint of its enormous investment in AI infrastructure and talent. The question is whether users will notice the difference, and whether the savings will translate into the kind of durable competitive advantage that justifies the risk of stepping back from the most powerful models on the market.
For now, the direction is clear. Under Suleyman, Microsoft is building toward a future in which Copilot runs increasingly on its own terms — and its own models.
---
Stay Ahead of AIGet the latest AI news, analysis, and breakthroughs — all in one place.
Read more AI news →





