On July 20, Microsoft confirmed an expanded partnership with AMD that will bring the chipmaker’s first rack-scale AI system, Helios, to Azure data centers. The deal also introduces three new Azure virtual machine families—HDv2, HXv2, and ND MI455X v7—powered by AMD’s upcoming “Venice” EPYC processors and Instinct MI455X GPUs. For enterprises and developers building AI on Azure, it means a credible alternative to Nvidia-based infrastructure is now on the roadmap, with the first Helios racks expected to ship for volume deployment in the second half of 2026.
What’s Actually Changing: Helios and Three New Azure VM Families
AMD’s Helios is not a single server but a rack-scale reference design that integrates up to 72 MI455X accelerators, 6th Gen EPYC CPUs, and Pensando networking into a liquid-cooled, double-wide rack. Each compute tray holds four GPUs and one CPU. AMD rates the full configuration at up to 2.9 exaFLOPS of FP4 performance, with 31 terabytes of HBM4 memory across the GPUs. The design follows Open Compute Project standards and uses open interconnects like UALink and Ultra Ethernet, aiming to give hyperscalers a customizable blueprint rather than a fixed appliance.
Microsoft will use Helios to power its new Azure ND MI455X v7 virtual machines, which are tailored for frontier-model inference, reasoning, and agent-oriented workloads. While exact per-instance GPU counts, pricing, and regional availability haven’t been disclosed, the announcement signals that production-scale inference on AMD hardware will be an option in Azure alongside Nvidia and Microsoft’s own Maia accelerators.
Beyond the GPU-focused instances, Microsoft also revealed two CPU-driven VM families built on AMD’s forthcoming EPYC “Venice” processors, with launches planned for 2026.
| VM Family | CPU Details | Memory / Storage | Networking | Target Workloads |
|---|---|---|---|---|
| Azure HDv2 | ~500 physical Venice cores | 4 TB RAM, 32 TB NVMe | 400 Gbps Azure Boost | Data prep, agent coordination, search |
| Azure HXv2 | 176 Venice cores (>5 GHz), 3D V-Cache | Up to 4 TB RAM | 800 Gbps InfiniBand | EDA, semiconductor design, technical simulation |
| ND MI455X v7 | Helios rack with MI455X GPUs + Venice CPU | 31 TB HBM4 per rack, per-instance GPU mem TBA | Pensando networking | Frontier-model inference, reasoning, agent AI |
Azure HDv2: Heavy Lifting for Data and Agents
The HDv2 series targets the often-overlooked work that surrounds AI models: data preparation, search, reinforcement learning pipelines, and agent coordination. Each VM will offer nearly 500 physical Venice cores, 4 TB of RAM, 32 TB of local NVMe storage, and 400 Gbps of Azure Boost networking. That makes it a strong fit for memory-intensive batch processing and orchestration of multi-agent systems—areas where raw GPU count isn’t the bottleneck.
Azure HXv2: High-Frequency Precision for Chip Design
HXv2 is designed for electronic design automation (EDA), technical computing, and large-scale simulations. With 176 Venice cores running above 5 GHz, 3D V-Cache, up to 4 TB of memory, and 800 Gbps InfiniBand for fast MPI communication, this family directly addresses workloads like semiconductor design, finite-element analysis, and computational fluid dynamics. Microsoft explicitly mentioned chip design as a target, reflecting the growing demand for cloud-based EDA tools.
ND MI455X v7: The Inference Powerhouse
The star of the announcement is the ND MI455X v7, the first Azure VM family to run on Helios racks. Microsoft says it is built for “reasoning, search, and agent-oriented inference workloads.” While details are scant, industry expectations peg this as a high-memory configuration suitable for large language models that need massive HBM4 pools—each Helios rack packs 31 TB of HBM4, or roughly 430 GB per GPU. For comparison, Nvidia’s current H100 offers 80 GB, and upcoming hardware pushes higher. AMD’s per-GPU memory advantage could be a differentiator for models that strain memory bandwidth.
What This Means for You: Developers, IT Pros, and Azure Users
The immediate impact depends on who you are.
For AI Developers and Data Scientists: If you’re building or fine-tuning models on Azure, the Helios-powered instances could eventually offer a cost-performance alternative to Nvidia’s Grace Blackwell systems. AMD asserts that Helios provides a lower cost per token for inference—a key metric for services that serve millions of queries. But the proof will be in real-world performance and software maturity. AMD’s ROCm stack supports major frameworks like PyTorch, TensorFlow, JAX, vLLM, and Triton, but developers will need to validate that their deployment pipelines, profiling tools, and model-serving frameworks translate smoothly to MI455X GPUs. Early testing will be crucial once preview instances become available.
For IT Administrators and Cloud Architects: The new CPU-based VM families (HDv2 and HXv2) are interesting for workloads that don’t require GPUs. HDv2’s 500 cores and massive memory could handle data orchestration and agent backends, while HXv2’s high clock speeds and InfiniBand target specialized HPC. Both families will likely run Linux-based operating systems, which aligns with Microsoft’s existing HPC portfolio. Windows admins may find indirect value if their organization uses Azure AI services that run on this infrastructure behind the scenes.
The heterogeneous strategy means you’ll have more procurement options: Nvidia, AMD, and Microsoft’s own chips. That could lead to better pricing leverage and reduced risk of supply constraints, though it also means more combinations to evaluate and manage.
For Regular Azure Customers and Windows Users: If you’re not directly managing cloud instances, you might still benefit. Cheaper or faster inference could lower the cost of Azure OpenAI Service, Copilot integrations, or other AI features. However, these downstream effects are speculative for now.
How We Got Here: AMD’s Long Game Against Nvidia’s Datacenter Dominance
AMD has been building its Instinct GPU line for years, but Nvidia’s CUDA ecosystem and early lead in AI training created a formidable moat. The Helios rack design is AMD’s answer to Nvidia’s DGX and NVL rack-scale systems, which bundle Grace CPUs and Blackwell GPUs into turnkey AI clusters. Microsoft’s choice is a landmark because Azure is one of the world’s largest consumers of Nvidia AI infrastructure, yet it is clearly hedging with multiple suppliers.
The foundations of this deal were laid with earlier AMD wins. In February, Meta announced a multi-year agreement to deploy up to 6 gigawatts of AMD Instinct GPUs, starting with Helios-based racks in the second half of 2026. OpenAI, Oracle, and TCS have also announced adoption of Helios-based systems. AMD claims eight of the top ten AI companies already use Instinct GPUs in some capacity.
Microsoft’s own hardware strategy is eclectic. It designs its own Maia AI accelerators and Cobalt CPUs but continues to buy aggressively from Nvidia. Adding AMD fits a pattern: offering customers choice while ensuring supply chain diversity. The partnership also spans CPUs, GPUs, networking hardware, and software, signaling a broad-based commitment.
What You Should Do Now
There’s no need to take immediate action—the Helios instances won’t arrive until at least the second half of 2026, and Microsoft hasn’t provided a firm date, region list, or pricing. However, if your organization relies on Azure for AI, consider these steps:
- Evaluate your inference workloads: Identify models that are memory-bound or sensitive to per-token cost. AMD’s HBM4 capacity could be a game-changer for large-parameter inference.
- Monitor ROCm development: Keep an eye on AMD’s software stack progress, especially support for your preferred frameworks and model-serving tools like vLLM.
- Plan for multi-vendor testing: When ND MI455X v7 instances reach preview, benchmark them against existing Nvidia options using your own queries and throughput requirements.
- Consider HDv2/HXv2 for non-GPU needs: If you run data preprocessing, EDA, or simulation workloads, the new Venice-based VMs may offer better price-performance than current offerings. Track announcements for early access programs.
- Engage with Microsoft account teams: Large customers can often get roadmaps and influence early availability. Express interest to your Azure representative.
What Comes Next
The July 20 announcement is a declaration of intent, not a launch. The real test will be when the first Helios racks light up in Azure data centers. AMD must prove that its open-standards rack design can match Nvidia’s vertically integrated systems in real-world inference throughput, reliability, and developer experience. Microsoft, meanwhile, must demonstrate that its orchestration and management layers can seamlessly support AMD GPUs alongside existing options.
For the broader Windows ecosystem, the impact is likely indirect but meaningful: a more competitive AI infrastructure market could eventually make Windows-based AI tools and services faster and cheaper. For now, mark your calendars for the second half of 2026—that’s when the rubber meets the road for AMD’s Helios in the cloud.