Microsoft quietly crossed a strategic threshold on Tuesday. For the first time, tens of thousands of AI prompts inside Excel and Outlook are being completed each week not by OpenAI or Anthropic, but by Microsoft’s own in-house MAI models — a shift reported by Bloomberg and confirmed by multiple secondary sources. The move is directional and intentional, not experimental: Microsoft AI CEO Mustafa Suleyman has named cost reduction as the explicit goal, and the architecture behind the company’s new models makes the math work in ways no simple “build vs. buy” framing captures.

For the hundreds of millions of people who use Office every day, the short-term experience will likely feel unchanged. For OpenAI and Anthropic — and for every enterprise team building on Azure AI — Tuesday’s disclosure marks the moment Microsoft’s in-house AI program moved from announcement to operational reality.

Microsoft Is Running Its Own AI Where It Used to Pay OpenAI and Anthropic

Excel and Outlook previously routed AI requests more heavily to models from OpenAI and Anthropic. Now, a portion of that volume runs on Microsoft’s internally built MAI family, according to a source familiar with the work who spoke to Bloomberg on condition of anonymity. A Microsoft spokesperson declined to comment.

The scale is still modest: tens of thousands of weekly prompts against a Copilot total that processes many millions. But the direction is the story. Microsoft is targeting the commodity layer of AI work — the drafting of email replies, the summarizing of threads, the generation of simple spreadsheet formulas — and routing those high-volume, low-complexity tasks to models it owns outright. Tasks requiring frontier-grade reasoning can still route to OpenAI or Anthropic. What is shifting is the enormous, repetitive inference load underneath those frontier calls, and that load is where third-party API bills actually accumulate.

How a 1-Trillion-Parameter Model Becomes Cheap Enough to Deploy in a Spreadsheet

MAI-Thinking-1, Microsoft’s flagship reasoning model launched at Build 2026 on June 2, carries approximately 1 trillion total parameters — making it one of the largest models ever built. Yet it activates only around 35 billion of those parameters per inference call, according to the MAI-Thinking-1 technical report.

The distinction is not semantic. MAI-Thinking-1 uses a sparse Mixture of Experts architecture: a gating network routes each incoming request to the specific subset of specialized sub-networks — “experts” — best suited to handle it. Only those activated experts consume compute during inference. The remaining ~965 billion parameters sit idle, consuming no power and incurring no per-call cost.

The economic consequence is what makes Microsoft’s cost argument credible. A model with 1 trillion total parameters can deliver reasoning quality comparable to frontier-class dense models while running at the inference cost of a 35-billion-parameter model. That is not a minor efficiency gain — it is the architectural mechanism behind the “10x cost efficiency” figure Suleyman cited in his Build 2026 keynote when Microsoft tuned a MAI model for McKinsey’s enterprise workloads and projected it outperforming GPT-5.5 on quality at roughly ten times lower cost, based on public pricing comparisons across model sizes.

MAI-Thinking-1 was trained on approximately 30 trillion tokens of commercially licensed data without distillation from any third-party model — meaning it does not borrow from GPT or Claude outputs, giving enterprise customers a clear data provenance chain. A 256,000-token context window lets it process roughly 600 pages of text in a single pass.

One critical caveat belongs here: Microsoft’s performance claims were largely produced by evaluations Microsoft commissioned. In blind side-by-side tests run by Surge, described as Microsoft’s independent human rating partner, evaluators across 1,276 tasks preferred MAI-Thinking-1 over Anthropic’s Claude Sonnet 4.6. On SWE-Bench Pro, a software engineering benchmark, Microsoft reports MAI-Thinking-1 scores 52.8%, which it says matches Claude Opus 4.6 on coding, per Microsoft’s benchmark results. Independent benchmark aggregator BenchLM.ai currently ranks MAI-Thinking-1 at No. 45 of 124 tracked models overall, with its strongest category being instruction following. AI researcher Andrej Karpathy described 2025 as an “evaluation crisis” in which standard benchmarks no longer reliably rank frontier models, in his 2025 year-in-review post — context that makes any self-commissioned evaluation worth treating as a signal, not a verdict.

Suleyman’s Explicit Goal: Eliminate Anthropic Spending

“Anthropic is extremely expensive and I think many people are urgently looking for alternatives,” Suleyman said in an interview last month. “We pay a lot of money to Anthropic, so our goal is to reduce and ultimately eliminate that cost.” The statement was made in June, roughly a week before Microsoft unveiled its full MAI lineup at Build 2026.

The financial context makes the timing matter. Anthropic projected roughly $10.9 billion in Q2 2026 revenue and about $559 million in operating income — what would be its first-ever profitable quarter, according to internal projections reported by CNBC and the Wall Street Journal. Those figures are projections, not audited results, and Anthropic has explicitly told investors it does not expect profitability to persist once full compute commitments take effect later in 2026. Still, the company filed a confidential S-1 with the SEC on June 1 at a post-money valuation of $965 billion — making Microsoft’s stated goal of “eliminating” its Anthropic spending a declaration of intent against one of its own portfolio companies on the eve of its IPO.

What Actually Changed About the OpenAI Partnership

Microsoft’s relationship with OpenAI underwent a material restructuring in April 2026. The renegotiated agreement ended Microsoft’s exclusive license to OpenAI’s intellectual property — meaning OpenAI can now sell through AWS, Google Cloud, and other cloud providers — while preserving a non-exclusive license that runs through 2032. Microsoft also removed its revenue-share obligation to OpenAI, with OpenAI retaining a capped revenue share arrangement with Microsoft through 2030.

William Blair analyst Jason Ader summarized the strategic logic: securing IP rights through 2032 protects the foundation of Microsoft’s Copilot strategy while freeing Azure to compete more actively for OpenAI workloads. NYU Stern professor Robert Seamans described it as Microsoft “continuing to rely on a really, really important partner, but also hedging their bets.”

With the 2032 deadline now a fixed calendar date rather than contingent on AGI declarations, Microsoft has a known expiration on its discounted access to OpenAI technology. Building MAI now is the hedge against paying full market rate when that license expires.

What This Means for Office Users — Including What Microsoft Isn’t Saying

For typical Copilot users, the immediate experience is unlikely to feel different. Microsoft is routing commodity tasks — inbox summaries, formula generation, simple chart creation — to MAI, not complex multi-step reasoning workflows. The company has consistently framed MAI deployment as targeting the tasks “good enough” AI handles well, not the tasks where frontier intelligence is genuinely required.

But there is a tension worth naming. Microsoft has not publicly disclosed which model completes which Copilot request, and it has not announced that the swap has occurred. Office users now interacting with AI-completed tasks in Excel or Outlook may be receiving MAI responses without knowing it. The Decoder’s analysis observed that the model transition “could mean paying the same amount for weaker AI so that Microsoft can lower its own costs” — noting that formal independent benchmark rankings for MAI-Thinking-1 show it trailing the frontier by a wider margin than the human-preference Surge evaluations suggested.

Microsoft’s stated roadmap extends the pattern: a Microsoft-built transcription model is expected to roll out to Teams in the coming months, and the company is building Copilot and Azure AI Foundry as multi-model platforms that route each prompt to the most cost-effective suitable option. One scenario Satya Nadella has gestured toward would make MAI models the default tier, with OpenAI or Anthropic models available as premium add-ons at additional cost to customers.

Where This Leaves OpenAI and Anthropic

The frontier labs are not losing partnership agreements. They are losing volume. The enormous repetitive inference load that accumulates at scale across hundreds of millions of daily Office interactions is exactly the tier MAI is designed to absorb. Losing that volume to Microsoft’s internal routing does not break any contract — it shows up as a revenue growth curve that grows more slowly than model adoption would otherwise suggest.

For Anthropic, the directness of Suleyman’s stated goal is unusual. Microsoft invested $5 billion in Anthropic and routes Claude to customers through Azure Foundry. It is simultaneously naming Anthropic cost elimination as a corporate objective. That tension will follow Anthropic into its IPO process, where Microsoft’s status as both major customer and declared cost-reduction target becomes a material disclosure.

For OpenAI, the calculus is more complex. Its discounted partnership with Microsoft remains intact through 2032 on model licensing. But OpenAI gained the right to sell through competitors — including AWS, which now offers OpenAI models alongside Anthropic’s Claude — meaning its distribution advantage through Microsoft is no longer exclusive. As Microsoft routes increasing volume to MAI, OpenAI’s consumption-based API revenue from Copilot workloads will compress even if the headline partnership survives.

The pattern is not unique to Microsoft. Google owns Gemini top-to-bottom with no licensing fees to any third party — and grew its cloud business at 63% year-over-year in the first quarter of 2026, compared to Azure’s 40%. Every major hyperscaler is running the same playbook: use third-party AI labs to build distribution, accumulate usage data and infrastructure, then build proprietary models for the high-volume commodity layer. The distributor has become the competitor. For the frontier labs, the commodity volume they counted on as a growth driver may be the first thing they lose.

Frequently Asked QuestionsDoes switching AI models inside Excel or Outlook change anything for users?

For most everyday tasks, probably not immediately. Microsoft is routing commodity requests — summarizing emails, generating spreadsheet formulas, drafting short replies — to its MAI models, while keeping more complex reasoning tasks on OpenAI or Anthropic. The practical quality difference for routine office work is likely minor. What users cannot currently do is see which model completed a given Copilot task. Separately, formal independent benchmark rankings for MAI-Thinking-1 show it trailing the frontier by a meaningful gap on some measures, even as Microsoft’s own human-preference evaluations suggest it competes well on everyday tasks.

What makes MAI-Thinking-1 capable enough for production use despite being described as “mid-sized”?

MAI-Thinking-1 uses a sparse Mixture of Experts architecture with approximately 1 trillion total parameters — only 35 billion of which activate per inference call. The “mid-sized” label refers to the active parameter count used per call, which determines inference cost. The total model capacity is among the largest ever built. The gating mechanism routes each request to the most relevant specialist sub-networks, achieving frontier-class reasoning quality at a fraction of the per-call compute cost of a dense model of comparable capability.

Is Anthropic’s revenue actually at risk from this?

The direct risk is to the volume of commodity inference Microsoft currently purchases from Anthropic at market rates. Anthropic projected roughly $10.9 billion in Q2 2026 revenue — internal projections shared with investors, not audited results — with Microsoft as one of its largest enterprise customers. Suleyman’s explicit goal to “eliminate” Anthropic spending targets exactly that relationship. The financial exposure is partly offset by Claude’s availability through AWS, Google Cloud, and direct API, but Microsoft’s distribution scale through 365 and Azure has been one of Anthropic’s most efficient customer-acquisition channels.

What happens when Microsoft’s OpenAI license expires in 2032?

The April 2026 partnership revision confirmed the license runs through 2032 as a fixed date. After that point, Microsoft would need to negotiate new terms or pay full market rates for OpenAI models. MAI is the hedge: by developing credible in-house models for commodity inference now, Microsoft reduces its exposure to whatever OpenAI charges in 2032. If MAI reaches frontier-class capability by then, Microsoft may not need to renew at all. That six-year runway is why Suleyman’s team is building with urgency.