AMD’s Helios Rack-Scale AI Servers Enter Full Production

The chipmaker takes direct aim at Nvidia’s data center dominance with a complete rack-scale system, already winning commitments from Microsoft, OpenAI and Anthropic.

AMD is making its most aggressive move yet to challenge Nvidia’s stranglehold on the AI data center market. The company announced on July 23 that its second-generation Helios rack-scale AI server has entered full production, with first customer shipments expected by the end of the third quarter of 2026.

The announcement came during AMD’s Advancing AI 2026 conference in San Francisco, where CEO Dr. Lisa Su detailed the company’s evolving AI infrastructure strategy. Helios represents a significant departure from AMD’s previous approach: rather than selling individual accelerators, the company is now offering a complete, pre-integrated rack system designed to compete head-to-head with Nvidia’s Vera Rubin NVL72 platform.

What’s inside the Helios rack

Each Helios rack integrates 72 AMD Instinct MI455X GPUs with 18 sixth-generation EPYC “Venice” CPUs, connected through AMD’s Pensando networking technology and accelerated by the ROCm open software stack. The MI455X is the first GPU built on AMD’s new CDNA 5 architecture, featuring a modular mix of 2nm and 3nm chiplets with 432GB of HBM4 memory and 23.3TB/s of peak memory bandwidth. A complete Helios system packs approximately 31 terabytes of combined HBM4 memory.

The Venice EPYC processor marks another milestone: it is AMD’s first 2nm server CPU, based on the new Zen 6 architecture, offering up to 256 cores and 512 threads per socket with PCIe 6.0 support. AMD claims Venice delivers up to 1.8 times the performance of the previous generation, calling it one of the largest generational leaps in EPYC history.

The competitive landscape

Helios is AMD’s answer to Nvidia’s Vera Rubin platform, which has dominated the emerging market for rack-scale AI systems. Where Nvidia has historically relied on proprietary interconnect technologies like NVLink, AMD is betting on an all-Ethernet, open-standards networking architecture. The company claims Helios delivers up to 30% more tokens per dollar than Nvidia’s Rubin NVL72.

“Helios is the highest-performing AI rack in the industry,” Su said during the keynote, adding that the system is designed “to train and run the world’s most demanding frontier models at massive scale”. She noted that frontrunner AI companies will deploy the system at gigawatt-scale.

A who’s who of AI customers

The customer list for Helios reads like a directory of the AI industry’s biggest players. Microsoft plans to deploy Helios extensively within its Azure cloud platform. OpenAI, Anthropic and Meta have all committed to gigawatt-scale deployments.

OpenAI has already been running GPT-class workloads on Helios racks for three months, according to Sachin Katti, the company’s head of infrastructure. “We expect that we’ll be deploying Helios at massive scale starting towards the end of this year and then accelerating towards 2027,” Katti said during the keynote. OpenAI holds warrants for up to 10% of AMD stock under a six-gigawatt agreement.

Anthropic announced a strategic partnership with AMD on the same day, covering up to two gigawatts of MI450-series GPUs plus up to $5 billion in AMD equity investment. AMD expects to begin shipping the first gigawatt in the first half of 2027. Tom Brown, Anthropic’s co-founder and chief compute officer, described how the lab evaluated AMD hardware: a single engineer connected a prior-generation MI355X rack to Claude, told the AI to bring up the machine, and left for the weekend. “We ended up with a graph of the actual performance of our leading model on it, just going up and up and up over the weekend,” Brown said.

Other adopters include Oracle, HUMAIN, TensorWave, Vultr and Cirrascale. The systems will be supplied by OEM partners including HPE, Lenovo, Supermicro and Bull, with infrastructure partners Sanmina and Taiwan’s Wiwynn.

Shift in AI workloads drives demand

Su framed the Helios launch within a broader shift in AI computing. She argued that the rise of agentic AI — autonomous AI systems that can execute multi-step tasks — is driving a “step-function increase in compute demand”. “When you ask an agent to perform a task, it actually involves dozens of steps. It has to reason, call tools, access data, and repeat this process until the problem is solved. So you need a lot of GPUs to support all of this,” Su explained.

She noted that 60% of computing resources are now used for inference rather than training. The company now projects the AI accelerator market will reach approximately $1.4 trillion by 2030 — a figure that would approach the size of today’s entire semiconductor industry. AMD also expects the server CPU market to exceed $200 billion by 2030.

Production and delivery timeline

Su clarified that first Helios shipments will begin in September, with production ramping through the fourth quarter and into the first half of 2027. “We’ve actually built the ramp this way because it is a complex system,” she said, adding that AMD wants original design manufacturers to fine-tune the manufacturing process and align shipments with customer data center buildouts.

Demand for the MI450 series is running above AMD’s expectations, Su said. The company has already planned capacity through 2027 and beyond.

For AMD, Helios is more than just another product launch. It represents a bet that the future of AI computing belongs to complete, integrated systems — and that customers are ready for an alternative to Nvidia’s proprietary ecosystem. Whether that bet pays off will become clearer when the first racks ship in just a few weeks.

Leave Comment

Your email address will not be published. Required fields are marked *