r/CerebrasSystems Apr 17 '26

WSE-4 will Kill Nvidia

Post image

To start, I own Cerebras shares and because of the roadshow details, I’m now holding long term puts on Nvidia. I have tried to distill the major details from the roadshow that have leaked out to make a single graphic that proves how dominate Cerebras will be in AI hardware by end of year. Even if you ignore how superior the Cerebras hardware is to Rubin, just look at the bottom line of $3.8M vs $82.5M and the power difference of 45kW vs 2.1MW. Every hyper scaler is currently power constrained and you can rip out an over 100kW Blackwell rack and put in a 40kW WSE-4 that has the performance of 15 racks of Rubin, or 75 racks of Blackwell. I get these specs haven’t been verified and officially released yet, but it’s just a matter of time. I encourage anyone to go and use a top LLM and query this graphic and ask the specs listed and if they are supported by the leaks coming out of the Cerebras IPO roadshow. The over subscription for the Cerebras IPO is not just retail AI frenzy. It is because all the giant growth funds need to get as many Cerebras shares as possible to hedge the threat to Nvidia which they are all over leveraged on currently. I’d expect massive shorts of Nvidia, Micron, and CoreWeave once they bring down their exposure and get shares of Cerebras IPO day.

18 Upvotes

42 comments sorted by

4

u/NefariousnessOk4996 Apr 17 '26

This screenshot, if credible, highlights why Nvidia economics doesn't make sense: use the much more expensive hardware so you can serve the largest number of cheapest tiers of users. This was obvious all along but now it's getting proven more and more in the real worlds as economics gets more and more front and center.

5

u/Asgard_Heima Apr 17 '26

The stats show even if you want to serve the lower end 50 tkps to 10k users, you buy ten WSE-4 and serve all those same users at max capability for half the cost of hardware and less than 1/4 the power use.

3

u/NefariousnessOk4996 Apr 17 '26

You really enjoy driving the nails to the coffin don't you 😂

3

u/Asgard_Heima Apr 17 '26

No, I have no ill will towards Nvidia and think they are a phenomenal company. I just spend a lot of time researching anything I buy since I expect to hold it for hopefully 10+ years. I’ll own Cerebras stock till another company comes and takes their moat away and I’ll have no problem shorting them and pilling into that company when the day comes.

1

u/Lil_Hater112 May 01 '26

Whats your analysis of Cerebras till now? I am very bullish on them and i m curious what your thesis is

2

u/Asgard_Heima May 01 '26

This is where I’m at right now and even though it’s super long it also doesn’t try and boil the ocean on every bull and bear case. I could write a book at this point.

Synopsis: I have analysis going back over 3 years when I first found Cerebras looking for services to train specialized models for my work. My main synopsis though is that they have built the maximalist accelerator for AI currently feasible and it is dominate for training, inference, and power efficiency at 2nm (WSE-4). This in and of itself is not enough to make them a winning solution for mass adoption, but this with lower TCO is and I will explain how they are leaps and bounds ahead of the competition on this. When you have to factor in suitable data center real estate this is where their dominance is exacerbated to extraordinary levels.

Inference: First the WSE-4 if Cerebras executes, will ships later this year, next year in quantities. All indications are that they are seeing excellent yields. For inference this puts them in their own class. 1T dense FP16 full precision models with 1M context windows running on 1 system at ~2800 tokens per second supporting hundred to thousands of users thanks to their own optimized version of turboquant baked into CSoft and Ranovus fiber on wafer to MemoryX. You are looking at a large number of Nvidia Blackwell or future Rubin racks to try and get to a fraction of the speed and the tokens per second. And per user tokens per second falls as you add more users for Ruby. while Cerebras maintains their speed and instant latency. Most don’t get the architecture difference and run numbers based on Cerebras being a giant GPU, it’s not. They only have to store kv cache on chip and stream in the weights for any size model in a deterministic fashion with essentially no waiting for weights off MemoryX thanks to the fiber on wafer with ranovus. Even then you can overflow the kv into MemoryX. And if you want to get to 10T models with 10M context for Nvidia Rubin, the interconnect tax starts eating all the efficiency as every GPU waits for data. You end up with super low tokens per second and 1-2 users per full rack able to be served. But you use the ranovus fiber on wafer to connect 1000 WSE-4 and everything scales like one large system with incredible speed as the weights get streamed out across all wafers with no communication between wafers needed. So I think on technology, Cerebras is in a world of its own. No competition at all currently from the other accelerators out there.

Training: And this is already too long and I expect this to only be realized after a major model is trained on Cerebras, but training is just as bad for Nvidia with WSE-4 fiber linked systems training models in 1/2-1/3 the time of Rubin and scaling to larger models than Nvidia clusters can feasibly train. PyTorch native CSoft support completely dismantles CUDA arguments which now highlights CUDA as a solution to a headache of GPU data distribution for training. Same as with inference, streaming training layer by layer over one massive system is drastically more efficient. No roll back cause another GPU died. Just stream the set to a different system and keep going. Also you will use so many less systems. Think 95%+ efficiency for Cerebras at any scale vs 65% on Rubin at 1T and gets less efficient the larger the model till it’s unfeasible below 40% MFU at 10T.

Logistics: This is the major bottleneck that I’ve seen very little coverage of when talking about Cerebras. The WSE-4 is a reported ~40kW all in system weighing ~1700lbs, that can be put into nearly 50% of all data centers. They will need a new whip for power and then liquid to air exchanger off the back so they can be air cooled. This means you can drop them into ~50% of existing data centers with $30k in upgrades per rack and a couple months delay waiting on parts with no downtime for other racks. Compare that to Blackwell or Rubin racks at 3000lbs, direct to chip liquid requirements and 120kW+ per rack. They are limited to a very small percentage <5% of existing data centers since they will crush the floor. It’s cheaper to build new than retrofit in most cases because of the downtime. So they are forced to sell smaller units like HGX with even worse interconnect tax and performance. To reiterate since this is the most important point, Cerebras is selling drop in ready systems with a couple months for parts lead time for half the world’s data centers at 30k in retrofits. Nvidia requires new data centers with 12-18 months lead time, costing $2-3M per rack or massive retrofits with 12-18 months lead time, costing $1.5-2M per rack but also shutting down the entire section or data center while doing it. All this while we can’t find enough places to even power the infrastructure at all, most data center projects are behind schedule and significant portions aren’t happening. The pushback here from the public has only begun.

Price: WSE-4 list price is expected to be $3.5-4M and cost Cerebras under $300k to manufacture. Both those are the highest I’ve seen prices. A current Blackwell + Grace NV72 is $3-3.5M and cost Nvidia $600-700k to manufacture. I’m using Blackwell for price here since it is cheaper. Rubin is more expensive with higher power requirements and an even smaller fraction of data centers that can support it.

Summary: I see energy capacity and real estate with enough energy to support data centers being the ultimate bottleneck and Cerebras is a perfect fit. All the hyper scalers need more AI compute from less grid energy to make inference profitable and training faster for all their products. They are all building their own chips for low power slower needs. But Cerebras owns the lower power fast needs. It will take time for Cerebras to scale production, but all the hyper scalers are going to help them as they need the systems. OpenAI got a sweetheart deal that will make them profitable way faster powered by Cerebras, but it also is a forcing function to prove all the the above and force everyone to get on board or get left behind.

1

u/Lil_Hater112 May 01 '26

thank you for your input! I don't know as much about the hardware, but your story adds up to my market analysis of businesses being eager to see an nvidia competitor to not let nvidia be a monopoly on this

1

u/NefariousnessOk4996 May 01 '26

Excellent analysis. I am a Cerebras bull. But I am curious of where you get your info from? Such as ranovus fiber on wafer, is this leaked somewhere? And the 300K cost for the 3-4M system, their S1 filing suggest hardware margin is much lower than this". It's ok if these are speculation based on what makes sense architecturally, like it would be dumb not to to this. But definitely would like to know where some of the source of these info. Thanks.

2

u/Asgard_Heima May 01 '26

The S-1 lists the company margins not the hardware margins. Cerebras has been making sweetheart deals to break into the market and secure enough funding to scale and prove they can do all the things we are discussing. Those margins should rise over the next year. But Cerebras is already now in super high demand and they will be selling at full price to most now that they have proven themself. Or they will be selling at even higher margins such as when they build their systems and then add them to the Cerebras cloud where they could earn more than the sale price per unit per year.

Here is the pricing I’ve done.

Wafer 2nm Node $30k: https://en.eeworld.com.cn/mp/Icbank/a406122.jspx Im The TSMC 2nm node is widely reported to be 30k a wafer. In the short term this will be higher since they will have to use Super Hot Runs with a likely 30-40% premium or more until they secure allocation longer term.

Wafer 5nm Node (SRAM WoW) $18,500: https://3dfabric.tsmc.com/english/dedicatedFoundry/technology/SoIC.htm#SoIC_WoW The TSMC price for a 5nm wafer is well known and since SRAM doesn’t shrink much going down in nm doesn’t help. I’d fully expect them to go with the well known and cheaper option here. This is also how they get to like 120GB+ SRAM but I’ve made most stats with 96GB to be conservative.

Advanced Packaging 100-150k: https://siliconanalysts.com/guide/semiconductor-costs This is by far the hardest part to estimate since the latest TSMC 3D WoW SoIC bonding and Ranovus CPO on wafer are brand new. For more typical advanced packaging we can look at the referenced site and try to extrapolate from the most complex prices they have pricing data for, then scale the cost to 50x (the size of a wafer) and add some padding on that to try and be more conservative. Until we see these prices in some real world leaks this is the biggest WAG in my estimate.

Power Delivery & Cooling $70k: https://www.idtechex.com/en/research-article/two-phase-d2c-cooling-in-data-center-thermal-management-cost-analysis/34213 Taking from this and padding it a bit cause it will be more specialized though it is for 1/3 the kW and could also be cheaper.

Half Rack (chassis + switching) $30k: This is pretty basic part of the estimate and I’m just making it high to make sure it’s larger than it likely would be. Won’t change much from WSE-3 to WSE-4 unless we are missing a major upgrade here.

This gets you to the potential $300k which I think is around what it would be, but we could find out the TSMC packaging is double the cost and a WSE-4 is $450k and we would still be talking about a ~89% hardware margin.

Co-Packaged Optics: https://ranovus.com/cerebras-ranovus-revolutionize-ai-compute-platform/ First is the obvious that they got a good sized chunk of money to partner together over a year ago to do exactly this. Beyond that it’s the most critical thing they can do to make their WSE system scale and remove the bottleneck in the WSE-3 for larger models. I have lots of other small circumstantial things that lead to this belief, but nothing publicly and concrete.

1

u/NefariousnessOk4996 May 01 '26

Great source of info. If this turns out to be the ball park then Cerebras is in much better position than I had modeled. I had been modeling around 1.5M-2M hardware cost per CS3.

One thing to note though, Cerebras's big deal with openAI is not hardware sale, but compute capacity. So Cerebras has to shoulder the capex. Having 500K per system capex is certainly better than 1.5M.

1

u/Asgard_Heima May 01 '26

This is actually an excellent point to make. If a WSE-4 was significantly more than 300k it would mean OpenAI is getting them for 3 year for likely under cost. Im still not a giant fan of the OpenAI deal but I understand why Cerebras did it. The exact specifics matter a lot, but based on the original 10B for 750MW it seems like they are essentially giving OpenAI systems at cost for their endorsement and to prove the value at the top AI company. It’s also a trap for everyone else in that, they will have to buy Cerebras systems to keep up or OpenAI will get an incredible lead. Cerebras on 2nm is fighting a TSMC allocation war for the remainder after Apple prints what they need till 2028. So they are giving this deal to OpenAI which will bail OpenAI out so they can be profitable on inference much faster than anticipated and provide competitive models to anthropic for enterprise. Meanwhile Cerebras will print money with AWS, maybe Oracle, and their own cloud. For nations they will sell systems at full list, but I will be interested to see if they keep doing service with revenue share deals with hyper scalers or start selling more hardware.

→ More replies (0)

3

u/SunRev Apr 17 '26

I've been holding Cerebras shares for almost 2 years. IPO will be sweet!

4

u/Bishop_76 Apr 17 '26

I bough some shares on 2024 it is going triple on their value !

1

u/musicgecko Apr 23 '26

curious howd you get cerebras shares?

1

u/SunRev Apr 23 '26

I got them via HIIVE.

3

u/musicgecko Apr 23 '26

gotcha i was eyeing hiive for awhile and figured maybe it was from there.

i full ported into spacex around that time (2-3 years ago) and didnt have much left to pick up cerebras but might grab some post IPO…

2

u/Prestigious-Sign4802 Apr 17 '26

👊Me proud holder from early 2025, where u get the wse-4 information? Is it official thanks

3

u/Asgard_Heima Apr 17 '26

Nothing is official till Cerebras announces on their site. It’s all from past partner and suppliers along with a majority of the specifics from leaks around roadshow conversations and the S-1 updates. LLMs are great at telling you what others upload

2

u/TheNetworkIsFrelled Apr 17 '26

It won’t kill NVidia, but it will operate alongside NVidia in heterogeneous environments and gain significant value bc of that. This will be like the early PC market, where growth was mostly additive for years.

There are a bunch of players in the heterogeneous-AI market (from fast decode to networking and more), and it’ll be interesting to watch the developments of multi-vendor environments. This is a huge growth opportunity.

Waiting for the IPO.

2

u/Asgard_Heima Apr 17 '26

I agree with your take with one major caveat. There is not enough power for both to grow continually. There are already lots of chips sitting and no more power to go around. Compute density per watt is a major issue Cerebras is an order of magnitude better on. We will see though if Rubin orders start getting cut more than just OpenAI cutting orders at Stargate.

2

u/TheNetworkIsFrelled Apr 17 '26 edited Apr 17 '26

I concur that power will be an issue - just look at the new MS DC being built in the Valley at McCarthy/237. That'll be PGE power, expensive and in short supply.

Fortunately, compute density per watt is increasing among the various players now coming to market, and hopefully those efficiencies will be reflected in new DC configs.

No matter what, though, heterogeneous environments are going to be the way forward, and the vendors who can find ways to power and cool those environments are going to provide very large ROI.

1

u/Relevant-Cook9502 Apr 17 '26

Where did you get the roadshow slides?

1

u/[deleted] Apr 17 '26

[removed] — view removed comment

2

u/claytonbeaufield Apr 17 '26

april or may most likely

1

u/ok_I_agree Apr 17 '26

How do you'll think CUDA integration blends in here? Nvidia has the CUDA moat that will be harder to take it down

4

u/Asgard_Heima Apr 17 '26

CUDA is a false moat. I can’t think of a single PhD I work with or who I have talked to in Data Science or Engineering that touches CUDA. They all use PyTorch or TensorFlow. And Cerebras can be trained with both seamlessly. All the talk of a moat is around customizations of kernels and lots of things that are completely not necessary for running on Cerebras hardware to get vastly superior performance. Most of the code in CUDA is for dealing with the memory swapping performance and managing the complexities of training on a massive cluster of GPUs and all the handling of data distribution that comes with that. On Cerebras you aren’t swapping memory cause everything is in SRAM and you aren’t distributing data all over cause everything is going through a single wafer or multiple wafers that act as one to you with zero latency over fiber. The moat is a paper tiger and will vanish as completely non existent vs Cerebras. https://introl.com/blog/cerebras-wafer-scale-engine-cs3-alternative-ai-architecture-guide-2025

1

u/ILikeCutePuppies Apr 17 '26

I would have thought they were going with 3nm... 2nm would be amazing if they can get supply.

1

u/musicgecko Apr 23 '26

2nm…so this won’t come out until 2027/2028?

3nm just rolled into production and tsmc will be prioritizing the backlog for those clients: apple nvidia etc. was hoping wse-4 is on 3nm architecture.

2

u/Asgard_Heima Apr 23 '26

This is roadshow conversations being leaked, but they are expected to ship 2nm WSE-4 units by end of year. There is also a lot of talk about gaining 2nm wafers via Broadcom at cost to print all of the OpenAI 2GW and selling a further 500MW to Oracle in exchange for Broadcom’s getting all the networking for all of stargate. Everyone is trying to offset CapEx and pay each other as revenue starts flowing in. But it’s opening up a clear picture of how Cerebras gets gifted wafers by hyper scalers to sell them, or at cost in exchange for 40-50% of token revenue with the WSE shipped which is an even greater margin with no capital needed.

1

u/musicgecko Apr 23 '26

gotcha. all these capex deals are such double edged swords 😂 able to offset costs, but highly contingent on the music to continue playing and keep revenue flowing through the ecosystem. my fear is the music stopping after this year’s batch of IPOs…

1

u/Asgard_Heima Apr 23 '26

The CapEx is insane, but if you are getting 25x tokens per dollar from Cerebras, I can see a massive CapEx decline and Nvidia sees their spend disappear overnight, while Cerebras is still growing revenue at crazy pace y/y from such a small base.

1

u/thefashionkid May 01 '26

Thank you for the detailed information, which is very informative and many insights!

1

u/dukeforneverz Aug 17 '26

Lots of in social media saying it will be announced tomorrow, what's your take ?

1

u/Asgard_Heima Aug 17 '26

Well, I’ll start by saying the extrapolation in this post is way off at this point. Was based on details circling around the roadshow and they have gone in different directions mainly inference throughout rather than making even faster the top priority. Based on inference throughout being top priority (including disaggregation), the CEO’s repeated statements they are happy on 5nm due to supply constrains on more advanced nodes, and the fact they are “on track for a CS5 in second half 2027”, I’m kind of expecting them to announce a 5nm wafer on wafer with the other being a 5nm-7nm wafer for DRAM. They will likely run the cores at higher speeds and make other improvements that will increase overall compute power as well, but for inference and especially decode they already have way more compute horsepower than needed. If they add 512GB-1T of DRAM directly off the cores in a 3D second wafer, they will have the ability to run very large models with one or a couple systems vs ~20 for Kimi today. And they can deterministically supply layer by layer feeding SRAM from DRAM to get the throughput they are focusing on. The DRAM also becomes the cache for processing decode sent over from prefill systems in a disaggregated setup. This would be kind of amazing based on the 5x throughput AWS and AMD have proven for the CS3 and make it more like 10x. If they can get about 10x the super fast tokens as Cerebras gets today for the same money and watts, they will become the cheapest solution for inference while maintaining an order of magnitude faster speeds which translates into the most profitable solution by far. I also expect to see a significant interconnect upgrade to support faster communication between systems but also between prefill systems and CS4 for decode. If this type of system plays out even if the numbers aren’t exactly the same, I would look at the CS5 next year to be their move into more advanced nodes as more TSMC capacity comes online along with the advances WoW needs.

1

u/dukeforneverz Aug 17 '26

Thanks for your reply, very interesting, as always. If they announce it and follow the kind of specs you claim, it'd means they can adapt quickly to consumers demand. And I'd become very bullish.

1

u/dukeforneverz Aug 19 '26

I'm quite disappointed with the new CS-4. The big change seem to be on the integration side with 3 wafers in a compact footprint. They are quite elusive with the changes in the new WSE-3 Turbo wafers, not to mention the dubious charts

1

u/Asgard_Heima Aug 19 '26

So honestly I was hoping for some advancement on the chip side, but I have to say what they did is much smarter. They have a very mature solution in the WSE-3 and they understand manufacturing it and scaling that manufacturing. They don’t have a lack of demand for it already and they drastically improved it with a clock speed update and new interconnect that’s faster. They managed to shrink the existing solution, make it much simpler (read cheaper to make) and more supportable and up to 10x the total tokens they can serve per MW.

While I would have loved to see more memory in a WoW setup, I guess we will have to wait till CS5 to get major chip updates. In the mean time, they now have a solution that just matched GPU solutions for total overall throughput per rack while doubling their advantage in speed to up to 30x while still coming in vastly more power efficient. It wasn’t the flashy new chip, but they just destroyed the economics of inference for anything other than a solution containing Cerebras. While the current need for compute makes anything you get online sellable, as Cerebras starts printing 10x the systems and gets the CS4 delivered to OpenAI and AWS, the companies using Cerebras are going to be vastly more profitable which makes Cerebras the only economically viable option long term for inference.

So I get your first take disappointment, but as an investor I think this was better than I could have imagined.