Microsoft will roll out AMD’s Helios rack-scale systems across its Azure cloud in the second half of 2026, adding a major new accelerator option for large-scale AI inference. The move introduces Azure ND MI455X v7 virtual machines built on AMD Instinct MI455X GPUs, plus two fresh CPU-only VM families—HDv2 and HXv2—that target data-heavy AI pipelines and electronic design automation.
What Microsoft and AMD actually announced
The centerpiece is the AMD Helios Rackscale Solution, a liquid-cooled design that packs 72 Instinct MI455X accelerators, sixth-gen EPYC “Venice” host processors, and Pensando networking into a single integrated rack. Microsoft confirmed it will deploy Helios to power Azure’s own AI services and customer-facing inference workloads, with volume availability beginning in the latter half of 2026.
Azure will get three new virtual machine families:
- ND MI455X v7 – Purpose-built for production inference on Helios, targeting reasoning, search, and agentic workloads that demand enormous pools of high-bandwidth memory.
- HDv2 – Large CPU instances with nearly 500 Venice cores, 4TB of RAM, 32TB of local NVMe storage, and 400Gb networking, designed for data prep, vector indexing, reinforcement learning, and agent coordination.
- HXv2 – 176-core, high-clock-speed VMs with up to 4TB of memory and 800Gb InfiniBand, tuned for chip simulation, scientific computing, and engineering analysis.
Microsoft also plans to broaden its use of AMD’s Pensando DPUs across selected Azure networking services and integrate AMD silicon more deeply with Azure Boost, the architecture that offloads virtualization, storage, and host management.
Why this matters for Azure customers
If you run AI workloads on Azure today, the Helios deployment introduces a credible alternative to Nvidia-backed instances. Here’s what that means in practice.
More capacity and pricing flexibility
The generative AI boom has made high-end GPU instances hard to get. Adding AMD-driven capacity could ease wait times for large deployments, particularly for inference rather than training. Customers who are not locked into Nvidia’s CUDA libraries could see lower per-token costs on AMD hardware, though Microsoft hasn’t published pricing yet.
A different performance profile
The MI455X accelerator carries up to 432GB of HBM4 memory—far more than most existing server GPUs. Across a 72-GPU rack, that totals roughly 31TB of high-bandwidth memory. For serving memory-hungry large language models with long context windows, that capacity can translate into fewer costly memory transfers and more consistent throughput. Real-world application performance will ultimately depend on software, but the spec sheet suggests AMD is betting on bandwidth as a differentiator.
Not a drop-in replacement for Nvidia
Moving an existing Nvidia workload to AMD is not a simple swap. Many AI pipelines have hidden CUDA dependencies—custom kernels, container images, monitoring agents, or serving frameworks that assume Nvidia hardware. Before migrating, teams should inventory their dependencies, run a representative benchmark suite on MI455X, validate model output quality, and plan a fallback path. The migration effort is manageable for many, but it won’t be zero-touch.
ROCm software readiness in the spotlight
AMD’s ROCm stack supports PyTorch, JAX, vLLM, and other major tools, but it lacks the years of fine-tuning that Nvidia’s CUDA ecosystem has accumulated. Microsoft is positioned to help: it can offer validated Azure container images, optimized model catalogs, and migration tooling. Early adopters should budget time for performance profiling and watch for Microsoft’s own guidance once preview instances appear.
What it means for everyday Windows users
You won’t rent a Helios rack yourself, but you will feel the downstream effects. Services you use—Copilot, GitHub Copilot, enterprise search, security analysis, and future agentic features—all run on Azure infrastructure. More cost-efficient inference could allow Microsoft to offer longer context windows, faster responses, or higher usage limits inside products you already pay for. Conversely, the company might use the savings to improve margins or fund more computationally intensive features without immediately changing the end-user price.
The other long-term thread is hybrid AI. Local neural processing units on Windows PCs handle routine, privacy-sensitive tasks, but the cloud still serves the largest models. A stronger, multi-vendor cloud back end makes it more likely that Windows features relying on heavy inference remain responsive and affordable.
How we got here
Microsoft has been building a deliberately heterogeneous cloud for years. It already runs AMD EPYC CPUs across a large slice of Azure, designs its own Maia AI accelerators, uses Arm-based Cobalt CPUs, and relies heavily on Nvidia GPUs. The Helios deal deepens the AMD partnership in a way that targets inference—the phase of AI that happens billions of times after a model is trained, and the one that dominates operational costs.
Nvidia’s dominance didn’t happen by accident. It built a tightly integrated platform of GPUs, interconnects, networking, software, and developer tools that make it hard for competitors to wedge in. AMD is responding with Helios, a blueprint that combines accelerators, CPUs, Pensando networking, liquid cooling, and open standards (Open Rack Wide, UALink, Ethernet) into a complete rack-scale design. By adopting it, Microsoft gets to test whether an open-standards approach can compete with Nvidia’s proprietary ecosystem without locking itself out of any single supplier’s roadmap.
What to do now if you’re an IT decision-maker
- Start the assessment, not the migration. Watch for Microsoft’s preview timeline and begin cataloguing your current AI workloads’ CUDA dependencies. Identify which models or inference services could realistically run on ROCm today.
- Build a benchmark suite. Collect representative prompts, batch sizes, context lengths, and concurrency levels. When AMD-powered VMs become available in early-access regions, you’ll want to run head-to-head comparisons on your own data, not just vendor-provided figures.
- Prepare for limited early availability. The first ND MI455X v7 capacity will likely be tight—Microsoft may prioritise its own services and strategic customers before broad on-demand access. Contact your Azure representative early if you’re interested.
- Consider data-movement workloads. The new HDv2 instances are aimed at AI data pipelines that often become hidden bottlenecks. If your team spends significant time on feature extraction, indexing, or batch preprocessing, those VMs could be a more immediate win than waiting for GPU availability.
- Watch for concrete pricing and SLA details. No public pricing or service-level commitments exist yet. Base infrastructure decisions on what’s published, not on launch-day announcements.
Outlook: what to watch next
The real test will come when independent benchmarks and customer case studies appear. Performance on paper—2.9 exaflops of FP4 per rack, 19.6TB/s of memory bandwidth per GPU—is only part of the story. The practical metrics are tokens per second per dollar, time-to-first-token under load, and how smoothly the ROCm tooling fits into existing MLOps pipelines.
Microsoft will also need to show it can scale. Deploying liquid-cooled, high-density racks demands power, upgraded data halls, and new operational procedures. The first regions to get Helios may be the ones already equipped for advanced cooling. Broad availability across numerous Azure geographies could take time.
Finally, keep an eye on the competition. Nvidia won’t stand still, and Microsoft’s own Maia accelerators add another variable. For now, the Helios deal gives Azure customers one more reason to watch the second half of 2026—and one more tool to keep AI infrastructure costs from slipping out of control.