TL;DR

Copilot Evaluation: Microsoft is weighing Kimi K3 for selected Copilot requests, but no deployment decision is confirmed. Cost Estimate: An unconfirmed estimate puts potential annual savings at up to $600 million, with workloads and assumptions undisclosed. Model Trade-off: Kimi K3 combines strong benchmark results with low API prices, while self-hosting and handoffs add costs and risk. Deployment Gates: Microsoft would need quality, safety, governance, and capacity tests before assigning production Copilot traffic.

Moonshot AI’s Kimi K3 model rankes near leading proprietary models, placing second on Vals AI’s index, third on Artificial Analysis’s Intelligence Index, and first in Frontend Code Arena. As a result, Microsoft is now weighing whether selected Copilot requests could move from OpenAI’s GPT and Anthropic’s Claude models to Kimi K3.

The Kimi K3 model could test whether benchmark strength translates to a large commercial assistant after confirmation through internal testing. Its possible Copilot role concerns only some tasks, not a wholesale provider change. A final decision may still be pending.

According to The Information, a partial move could reduce annual AI inference costs by up to $600 million, but the eligible workloads and assumptions remain undisclosed. Microsoft said paid Microsoft 365 Copilot seats exceeded 20 million in fiscal Q3 2026. Small per-answer differences can accumulate across millions of tasks, making lower-cost model options financially material at that scale.

Inference is the computing used each time an AI model answers a prompt. Model routing could reserve expensive systems for harder work while sending suitable prompts to cheaper options. Quality, latency, capacity, security, and policy tests would all precede a deployment choice.

How Model Routing Changes the Cost Equation

Microsoft’s Azure platform for selecting and deploying AI models already has a model router that balances performance and compute cost. Its Balanced, Quality, and Cost modes can restrict routing to a chosen set of models. Kimi K3 remains a candidate rather than becoming the main model used for Copilot.

Kimi K3 has 2.8 trillion parameters and uses a mixture-of-experts architecture, which activates only part of the model for each request. Specifically, 16 of 896 specialist components run at a time, while native vision and a one-million-token context window support larger jobs. Moonshot AI plans to release the full trained weights on July 27, after which they could be downloaded and self-hosted.

Launch pricing is $0.30 per million cache-hit input tokens, $3 per million cache-miss input tokens, and $15 per million output tokens. Open-weight models can alter the serving-cost calculation as they can be run at a fraction of the cost that OpenAI or Anthropic charge their clients. 

Kimi K3’s model size however does not guarantee quality as  can become unstable if reasoning history is lost when one model transfers a conversation to another. This would be a direct concern for routed Copilot conversations.

Microsoft would need to send comparable coding and reasoning prompts to Kimi K3, GPT, and Claude, then measure accuracy, latency, retry rates, reliability, and safety before assigning any workload. Validating the savings estimate would also require per-request Copilot volumes, hosting costs, and pass/fail thresholds. Weaker answers or more retries could consume extra computing resources, while dedicated hardware would add another expense.

Capacity Sets the Next Tests

Since March 2026, Copilot Cowork’s multi-model approach uses AI models from both Anthropic and OpenAI. Cowork separates generation from review: GPT drafts research responses while Claude checks accuracy, completeness, and citation quality. Another provider could offer more routing and pricing options, but each model must preserve conversation state and satisfy enterprise controls.

Meanwhile, Alibaba is preparing its Qwen 3.8 model for public release as another downloadable alternative to OpenAI and Anthropic. Supplier choice can improve Microsoft’s negotiating and routing options, although benchmark rankings and list prices cannot replace tests on actual Copilot tasks.

Possible procurement rules and security restrictions for Chinese AI models could affect whether Microsoft can deploy Kimi K3. Policy discussions had not produced a binding rule.

Moving from API fees to self-hosting would shift expenses toward hardware, energy, operations, security, and capacity planning. Moonshot AI said on July 21 that Kimi K3 demand during the preceding 48 hours nearly reached available compute and pressured its graphics processors. Kimi K3 must absorb sustained Copilot-scale demand without exhausting that capacity before Microsoft assigns production traffic.