AI

Microsoft Tests Moonshot AI's Kimi K3 for Copilot Workloads

Microsoft is internally testing Moonshot AI's Kimi K3 to handle part of Copilot's inference load, with potential annual cloud savings of up to $600 million.

Microsoft is internally testing Moonshot AI's Kimi K3 for Copilot, with internal estimates suggesting up to $600 million in annual cloud infrastructure savings. The evaluation covers coding, reasoning, and cost per token on specific workloads, and a deployment would face technical, security, and export-control reviews. The move signals Microsoft is preparing to route Copilot traffic across multiple model providers rather than rely primarily on OpenAI and Anthropic.

The math behind the move

Microsoft is internally evaluating Moonshot AI's Kimi K3 to handle a portion of Copilot's inference workload, according to a report from The Information. The goal is to reduce the share of requests currently routed to OpenAI and Anthropic models, and to lower the cloud bill that comes with serving Copilot at scale.

Microsoft has told internal teams that moving even part of Copilot traffic onto Kimi K3 could reduce annual cloud infrastructure costs by up to $600 million. The company has not committed to a final deployment, and any switch would still have to clear technical, security, and export-control reviews. Initial workloads, if the test is approved, would likely come from the less sensitive end of the Copilot product line.

Kimi K3 is being assessed on more than raw benchmark performance. Microsoft is also measuring cost per token on specific Copilot tasks, particularly coding and reasoning jobs, where Moonshot's pricing has been aggressive enough to attract enterprise interest outside China.

Why Microsoft needs a third rail

Copilot's growth has made Microsoft unusually dependent on a small set of suppliers. The current stack leans on OpenAI for the highest-end reasoning, on Anthropic for certain safety-sensitive workloads, and on Microsoft's own models for everything else. That concentration looks fine on a slide, but it is also a single point of failure in pricing negotiations, in rate-limit disputes, and in any future scenario where one of the suppliers changes terms.

Kimi K3 is a useful third option for a few reasons. It is strong enough on coding benchmarks to be a credible drop-in for parts of GitHub Copilot. It is cheap enough per token to absorb the long tail of lighter queries. And, perhaps most importantly, it gives Microsoft a politically useful hedge: a model that comes from a Chinese AI lab with serious commercial deployment experience, which complicates the geopolitical framing of any future US-versus-China AI export debate.

What is still unresolved

The Information's report is careful to note that no final decision has been made. The technical integration alone is non-trivial: routing policies, prompt formatting, evaluation harnesses, and the specific mix of tasks that move over all have to be rebuilt. Data-sovereignty questions around a Chinese model serving enterprise customers in Europe and North America are also still live.

Still, the direction of travel is clear. The era of Copilot as a single-vendor product is ending, and the era of Copilot as a multi-model router that picks the cheapest capable model per query has begun. Microsoft just gave itself a way to do that without giving OpenAI or Anthropic any more leverage over its pricing.