On July 23, AMD unveiled Helios, a rackscale AI infrastructure design that connects 72 Instinct MI455X GPUs into a single coherent compute domain for the first time. The move marks AMD’s most direct challenge to Nvidia’s data center dominance, shifting the company from selling individual accelerators to delivering entire, exaFLOP-class AI systems designed to compete rack-for-rack with Nvidia’s Vera Rubin NVL72 platform.
Inside the Helios Rack: What AMD Actually Announced
Helios is not a finished appliance but a reference architecture built around open standards. At its core sit 72 AMD Instinct MI455X accelerators, each packing 432GB of HBM4 memory and 23.3 TB/s of peak memory bandwidth. The GPUs are arranged in four-accelerator compute trays and linked through a UALink-over-Ethernet fabric, with AMD specifying up to 260TB/s of aggregate scale-up bandwidth inside the rack.
The system’s headline numbers are massive: up to 2.9 exaFLOPS of FP4 compute, 31TB of total HBM4 memory, and 43TB/s of scale-out bandwidth via Pensando Vulcano AI NICs. These figures represent AMD’s first coherent 72-GPU design — a critical step, because large AI models require accelerators to work together as a single logical resource, not as isolated islands.
AMD pairs the MI455X GPUs with the new EPYC 9006 Series “Venice” processors, based on Zen 6 cores. The CPUs handle AI host duties and agentic workloads, while the MI455X brings a revamped CDNA 5 architecture. The chiplet design uses TSMC’s advanced 2N process for the compute dies and 3N for fabric and I/O dies, along with a switch to 32-wide wavefronts for better instruction latency and branch divergence handling. The memory hierarchy also gets a major overhaul, trading the previous Infinity Cache for a higher-bandwidth shared L2 cache and doubling local data store capacity to 320KB per workgroup processor.
Why 31TB of HBM4 Could Change the Game
Memory capacity is Helios’s most tangible differentiator. AMD claims each MI455X offers 432GB — 50% more than Nvidia Vera Rubin’s 288GB per GPU — and that stacks up to 31.1TB across 72 accelerators, versus about 20.7TB for a comparable Nvidia NVL72 rack. For frontier models, long-context reasoning, and retrieval-heavy applications, local memory size directly determines how large a model partition can be and how much data movement is required. More HBM means fewer sharding headaches, better handling of longer context windows, and less time spent shuffling weights across the fabric.
That advantage could prove decisive in inference workloads, where data locality and memory bandwidth often matter more than raw floating-point throughput. AMD is betting that customers will see the extra memory as a practical advantage, even if peak FLOPS comparisons are more nuanced.
The Performance Claims: Sorting Hype from Hardware
AMD’s comparison table shows the MI455X besting Nvidia Rubin in several peak theoretical metrics: 40.26 PFLOPS OCP MXFP4 versus 35 PFLOPS NVFP4, 20.13 PFLOPS OCP MXFP8 versus 17.5 PFLOPS, and the aforementioned memory edge. The company also asserts up to 15% more OCP MXFP4 performance and as much as 30% more tokens per dollar in internal testing.
But these are vendor projections, and the details matter. OCP MXFP4 and Nvidia’s NVFP4 are not identical formats; direct comparisons require careful tuning and real-world validation. AMD itself acknowledges that realized performance will be lower than theoretical peaks, and the software ecosystem remains a critical variable. For enterprise buyers, the right questions are not just about specs but about delivered throughput for a specific model, context length, and latency budget — and how much engineering effort is needed to get there.
What This Means for Enterprise IT and Windows-Centric Organizations
Helios won’t appear in a retail box, but its influence will ripple through the services and clouds that Windows users depend on. For enterprise IT leaders, a credible AMD rackscale platform introduces several practical possibilities:
- More cloud choice: Microsoft has announced plans to deploy Helios at scale on Azure, meaning future Azure AI instances could offer AMD-powered options alongside Nvidia-based ones. This could let organizations pick the best cost-performance for their specific workloads.
- Potentially lower inference costs: If AMD’s tokens-per-dollar claims survive independent testing, businesses running large language models, retrieval-augmented generation, or document intelligence on Azure could see reduced bills.
- Reduced vendor concentration risk: A second large-scale GPU platform eases supply constraints and gives IT buyers negotiating leverage, especially as AI infrastructure capacity remains strategically scarce.
- Open hardware paths: Helios’s reference-design approach, built on Open Compute Project and Ultra Ethernet Consortium standards, may appeal to organizations that prefer less vertically locked-in infrastructure.
- High memory for long contexts: The 31TB HBM4 pool is tailor-made for increasingly long-context models and agentic workflows that Windows shops are beginning to adopt through Microsoft 365 Copilot and custom AI solutions.
The catch is software portability. Organizations heavily invested in CUDA-optimized code, proprietary kernels, and Nvidia-tuned deployment pipelines won’t switch overnight. Migration costs, staff retraining, and performance tuning are real friction points that hardware advantages alone can’t erase.
How We Got Here: AMD’s Journey from GPU Supplier to System Architect
For years, AMD’s Instinct line competed primarily at the accelerator level, offering compelling silicon but rarely matching Nvidia’s tightly integrated rack-scale designs. Nvidia’s success has always been as much about the system — combining GPUs, CPUs, networking, and software into a single AI factory building block — as about the GPU itself.
Helios changes that equation. It’s AMD’s first coherent 72-GPU rack design, a direct response to Nvidia’s NVL72 platforms used by both the current Blackwell and upcoming Vera Rubin generations. The inclusion of EPYC Venice CPUs, Pensando networking, and a reference architecture built around open standards signals that AMD now wants to be a system architect, not just a chip supplier.
This shift has been building. The MI400 series roadmap had long hinted at a rack-scale play, and AMD’s partnership with Cerebras — which offloads certain inference stages to specialized hardware — shows the company is thinking about AI infrastructure as a flexible, heterogeneous fabric rather than a monolithic GPU cluster.
What Should IT Decision-Makers Do Now?
Helios won’t be broadly available until later in 2026, but the planning cycle for AI infrastructure moves slowly. Here are four concrete steps:
- Audit your AI workload profiles. Determine whether your models are limited by HBM capacity, network bandwidth, or raw compute. If memory footprint or context length is your bottleneck, Helios deserves a closer look.
- Watch for independent benchmarks. Don’t rely on vendor peak figures. Wait for third-party testing that runs your model class (LLMs, diffusion models, retrieval pipelines) at the scale you need.
- Evaluate software ecosystem readiness. Check whether your existing AI frameworks and deployment pipelines are supported on AMD’s ROCm stack, and weigh the cost of porting any custom CUDA code against the potential hardware savings.
- Ask your cloud provider about timelines. If you’re an Azure customer, inquire about Helios instances’ expected availability, pricing, and supported services. Early engagement can also give you a voice in shaping the offering.
For most Windows-focused organizations, the immediate action is to stay informed and include AMD-based instances in future proof-of-concept tests. You don’t need to rip out existing infrastructure, but ignoring a viable second source would be shortsighted.
The Road Ahead: Partnerships and Pitfalls
AMD is securing strategic customers. A partnership with Anthropic targets up to 2 gigawatts of MI450-series GPU capacity, and Microsoft’s Azure deployment — though size is undisclosed — adds cloud credibility. The Cerebras collaboration, expected to launch through Cerebras Cloud in H2 2026, promises up to 5x better tokens per second per watt by combining AMD’s prefill hardware with Cerebras’s decode-specialized silicon.
Still, execution will determine everything. The AI infrastructure market rewards those who can deliver in volume, with dependable software, and at scale. Helios gives AMD its best platform yet to challenge Nvidia, but the real test begins when the first racks ship and buyers start measuring results in production environments. If AMD can turn impressive specs into reliable, cost-effective deployments, enterprise IT will be the ultimate winner.