Microsoft introduces Maia 200, a new AI inference chip aimed at cheaper, faster token generation
Microsoft unveiled Maia 200, a second-generation AI accelerator designed for inference at cloud scale. The company says the chip improves performance-per-dollar and is being deployed in Azure, reflecting the intensifying race among hyperscalers to reduce dependence on third-party GPUs.
- PUBLISHED
- UPDATED

Microsoft on Monday announced Maia 200, its next-generation AI accelerator built specifically to improve the cost and speed of running AI models in production. The company framed the chip as an inference-first design intended to make “token generation” cheaper at scale—an increasingly central metric as AI features spread through consumer and enterprise software.

According to Microsoft, Maia 200 is fabricated on TSMC’s 3-nanometer process and pairs low-precision compute (FP8/FP4) with a redesigned memory system to keep large models fed efficiently. Microsoft also emphasized system-level design choices, including a scale-up networking approach that relies on standard Ethernet rather than proprietary fabrics, arguing that this can deliver cost and reliability benefits in large clusters.
The competitive context is clear: hyperscalers want more control over the economics of AI serving. Nvidia remains dominant, but cloud providers increasingly pursue custom silicon so they can optimize hardware, networking, software tooling, and datacenter operations as one integrated stack. Microsoft positioned Maia 200 as its strongest “first-party” accelerator to date, with deployments already underway in an Azure U.S. region and more to follow.
Microsoft also highlighted developer enablement through a Maia SDK preview, aiming to reduce friction for teams optimizing models and kernels across heterogeneous hardware. If the rollout scales as planned, Maia 200 could become a meaningful lever for Azure’s AI margins—especially as inference demand grows faster than many organizations’ willingness to pay for it.