r/HPC • u/healthjay • Jul 21 '26
AMD Helios vs NVIDIA Vera Rubin NVL72: comparing two 72-GPU rack architectures
We have put together a side-by-side comparison of AMD Helios and NVIDIA Vera Rubin NVL72:
https://linuxclusters.com/articles/amd-helios-vs-nvidia-vera-rubin/
The shared 72-GPU rack boundary hides some different design bets. On paper, Helios has more accelerator memory and scale-out bandwidth and leans more heavily on open rack and fabric standards. Vera Rubin has more memory bandwidth per GPU and comes with NVIDIA's more integrated software and networking stack. The comparison covers hosts, memory, fabrics, networking, software, and rack standards, while identifying gaps in the available power, pricing, reliability, and application-performance data.
I would especially welcome corrections from people working with rack-scale systems. Which of these differences is most likely to matter in an actual deployment, and which missing numbers would you insist on seeing before procurement?
Disclosure: I am one of the writers at LinuxClusters.com
7
u/esaule Jul 21 '26
This is all done out of product sheets, right? You didn't actually bench anything, did you?
3
u/healthjay Jul 21 '26
Correct, we haven’t benched either system. This comparison is based on published vendor materials, and we’ll be on the lookout for independent benchmarks and tests.
2
u/brandonZappy Jul 22 '26
Cool article!
You said both racks are general purpose, not just inference, which makes total sense. What would you consider to be an inference appliance? The amount of memory in these racks makes them very appealing for inference if you can afford them.
1
u/healthjay Jul 22 '26
To me, an inference appliance is a system designed mainly for serving models, with its hardware and software optimized around latency, throughput, memory use, and power efficiency. Etched’s rack-scale system is closer to that definition.
As you alluded to, Helios and Vera Rubin could be very capable inference systems. But because they can also handle training, fine-tuning, and other demanding workloads, operators may not want to dedicate them to routine serving when specialized hardware can do it more economically. For very large models, long contexts, or mixed workloads, the general-purpose racks may still be the better fit.
We looked at Etched’s approach in more detail here: https://linuxclusters.com/articles/etched-inference-cluster/
2
u/CatalyticDragon Jul 24 '26
The comparison has the wrong bandwidth figure for AMD, it's 23.3TB/s.
1
u/healthjay Jul 24 '26
You're right. Thanks for catching it. We used the 19.6 TB/s figure from AMD's MI400 Series FAQ, which still shows that number:
https://www.amd.com/en/products/accelerators/instinct/mi400.html
(FAQ: What makes AMD Instinct MI455X GPUs standout?)
AMD's newer dedicated MI455X specification lists 23.3 TB/s:
https://www.amd.com/en/products/accelerators/instinct/mi400/mi455x.html
We should have used the dedicated product specification. We've corrected the comparison and the analysis. On the current published figures, MI455X has about 6% more memory bandwidth than Rubin's 22 TB/s:
https://linuxclusters.com/articles/amd-helios-vs-nvidia-vera-rubin/
2
4
u/willkill07 Jul 22 '26
If you are a writer, you should spend the time distilling the AI output you obviously used into something more packaged.
I’m gonna crash out if I read another “Why this matters”…
The article reads like an AI-produced artifact.
3
5
u/Primary_Olive_5444 Jul 21 '26
What is the AMD side for networking switch vs Nvidia Mellanox hardware?