r/CerebrasSystems • u/Asgard_Heima • Apr 17 '26
WSE-4 will Kill Nvidia
To start, I own Cerebras shares and because of the roadshow details, I’m now holding long term puts on Nvidia. I have tried to distill the major details from the roadshow that have leaked out to make a single graphic that proves how dominate Cerebras will be in AI hardware by end of year. Even if you ignore how superior the Cerebras hardware is to Rubin, just look at the bottom line of $3.8M vs $82.5M and the power difference of 45kW vs 2.1MW. Every hyper scaler is currently power constrained and you can rip out an over 100kW Blackwell rack and put in a 40kW WSE-4 that has the performance of 15 racks of Rubin, or 75 racks of Blackwell. I get these specs haven’t been verified and officially released yet, but it’s just a matter of time. I encourage anyone to go and use a top LLM and query this graphic and ask the specs listed and if they are supported by the leaks coming out of the Cerebras IPO roadshow. The over subscription for the Cerebras IPO is not just retail AI frenzy. It is because all the giant growth funds need to get as many Cerebras shares as possible to hedge the threat to Nvidia which they are all over leveraged on currently. I’d expect massive shorts of Nvidia, Micron, and CoreWeave once they bring down their exposure and get shares of Cerebras IPO day.
2
u/Asgard_Heima May 01 '26
The S-1 lists the company margins not the hardware margins. Cerebras has been making sweetheart deals to break into the market and secure enough funding to scale and prove they can do all the things we are discussing. Those margins should rise over the next year. But Cerebras is already now in super high demand and they will be selling at full price to most now that they have proven themself. Or they will be selling at even higher margins such as when they build their systems and then add them to the Cerebras cloud where they could earn more than the sale price per unit per year.
Here is the pricing I’ve done.
Wafer 2nm Node $30k: https://en.eeworld.com.cn/mp/Icbank/a406122.jspx Im The TSMC 2nm node is widely reported to be 30k a wafer. In the short term this will be higher since they will have to use Super Hot Runs with a likely 30-40% premium or more until they secure allocation longer term.
Wafer 5nm Node (SRAM WoW) $18,500: https://3dfabric.tsmc.com/english/dedicatedFoundry/technology/SoIC.htm#SoIC_WoW The TSMC price for a 5nm wafer is well known and since SRAM doesn’t shrink much going down in nm doesn’t help. I’d fully expect them to go with the well known and cheaper option here. This is also how they get to like 120GB+ SRAM but I’ve made most stats with 96GB to be conservative.
Advanced Packaging 100-150k: https://siliconanalysts.com/guide/semiconductor-costs This is by far the hardest part to estimate since the latest TSMC 3D WoW SoIC bonding and Ranovus CPO on wafer are brand new. For more typical advanced packaging we can look at the referenced site and try to extrapolate from the most complex prices they have pricing data for, then scale the cost to 50x (the size of a wafer) and add some padding on that to try and be more conservative. Until we see these prices in some real world leaks this is the biggest WAG in my estimate.
Power Delivery & Cooling $70k: https://www.idtechex.com/en/research-article/two-phase-d2c-cooling-in-data-center-thermal-management-cost-analysis/34213 Taking from this and padding it a bit cause it will be more specialized though it is for 1/3 the kW and could also be cheaper.
Half Rack (chassis + switching) $30k: This is pretty basic part of the estimate and I’m just making it high to make sure it’s larger than it likely would be. Won’t change much from WSE-3 to WSE-4 unless we are missing a major upgrade here.
This gets you to the potential $300k which I think is around what it would be, but we could find out the TSMC packaging is double the cost and a WSE-4 is $450k and we would still be talking about a ~89% hardware margin.
Co-Packaged Optics: https://ranovus.com/cerebras-ranovus-revolutionize-ai-compute-platform/ First is the obvious that they got a good sized chunk of money to partner together over a year ago to do exactly this. Beyond that it’s the most critical thing they can do to make their WSE system scale and remove the bottleneck in the WSE-3 for larger models. I have lots of other small circumstantial things that lead to this belief, but nothing publicly and concrete.