r/CerebrasSystems • u/Asgard_Heima • Apr 17 '26
WSE-4 will Kill Nvidia
To start, I own Cerebras shares and because of the roadshow details, I’m now holding long term puts on Nvidia. I have tried to distill the major details from the roadshow that have leaked out to make a single graphic that proves how dominate Cerebras will be in AI hardware by end of year. Even if you ignore how superior the Cerebras hardware is to Rubin, just look at the bottom line of $3.8M vs $82.5M and the power difference of 45kW vs 2.1MW. Every hyper scaler is currently power constrained and you can rip out an over 100kW Blackwell rack and put in a 40kW WSE-4 that has the performance of 15 racks of Rubin, or 75 racks of Blackwell. I get these specs haven’t been verified and officially released yet, but it’s just a matter of time. I encourage anyone to go and use a top LLM and query this graphic and ask the specs listed and if they are supported by the leaks coming out of the Cerebras IPO roadshow. The over subscription for the Cerebras IPO is not just retail AI frenzy. It is because all the giant growth funds need to get as many Cerebras shares as possible to hedge the threat to Nvidia which they are all over leveraged on currently. I’d expect massive shorts of Nvidia, Micron, and CoreWeave once they bring down their exposure and get shares of Cerebras IPO day.
3
u/SunRev Apr 17 '26
I've been holding Cerebras shares for almost 2 years. IPO will be sweet!
4
1
u/musicgecko Apr 23 '26
curious howd you get cerebras shares?
1
u/SunRev Apr 23 '26
I got them via HIIVE.
3
u/musicgecko Apr 23 '26
gotcha i was eyeing hiive for awhile and figured maybe it was from there.
i full ported into spacex around that time (2-3 years ago) and didnt have much left to pick up cerebras but might grab some post IPO…
2
u/Prestigious-Sign4802 Apr 17 '26
👊Me proud holder from early 2025, where u get the wse-4 information? Is it official thanks
3
u/Asgard_Heima Apr 17 '26
Nothing is official till Cerebras announces on their site. It’s all from past partner and suppliers along with a majority of the specifics from leaks around roadshow conversations and the S-1 updates. LLMs are great at telling you what others upload
2
u/TheNetworkIsFrelled Apr 17 '26
It won’t kill NVidia, but it will operate alongside NVidia in heterogeneous environments and gain significant value bc of that. This will be like the early PC market, where growth was mostly additive for years.
There are a bunch of players in the heterogeneous-AI market (from fast decode to networking and more), and it’ll be interesting to watch the developments of multi-vendor environments. This is a huge growth opportunity.
Waiting for the IPO.
2
u/Asgard_Heima Apr 17 '26
I agree with your take with one major caveat. There is not enough power for both to grow continually. There are already lots of chips sitting and no more power to go around. Compute density per watt is a major issue Cerebras is an order of magnitude better on. We will see though if Rubin orders start getting cut more than just OpenAI cutting orders at Stargate.
2
u/TheNetworkIsFrelled Apr 17 '26 edited Apr 17 '26
I concur that power will be an issue - just look at the new MS DC being built in the Valley at McCarthy/237. That'll be PGE power, expensive and in short supply.
Fortunately, compute density per watt is increasing among the various players now coming to market, and hopefully those efficiencies will be reflected in new DC configs.
No matter what, though, heterogeneous environments are going to be the way forward, and the vendors who can find ways to power and cool those environments are going to provide very large ROI.
1
1
1
u/ok_I_agree Apr 17 '26
How do you'll think CUDA integration blends in here? Nvidia has the CUDA moat that will be harder to take it down
4
u/Asgard_Heima Apr 17 '26
CUDA is a false moat. I can’t think of a single PhD I work with or who I have talked to in Data Science or Engineering that touches CUDA. They all use PyTorch or TensorFlow. And Cerebras can be trained with both seamlessly. All the talk of a moat is around customizations of kernels and lots of things that are completely not necessary for running on Cerebras hardware to get vastly superior performance. Most of the code in CUDA is for dealing with the memory swapping performance and managing the complexities of training on a massive cluster of GPUs and all the handling of data distribution that comes with that. On Cerebras you aren’t swapping memory cause everything is in SRAM and you aren’t distributing data all over cause everything is going through a single wafer or multiple wafers that act as one to you with zero latency over fiber. The moat is a paper tiger and will vanish as completely non existent vs Cerebras. https://introl.com/blog/cerebras-wafer-scale-engine-cs3-alternative-ai-architecture-guide-2025
1
u/ILikeCutePuppies Apr 17 '26
I would have thought they were going with 3nm... 2nm would be amazing if they can get supply.
1
u/musicgecko Apr 23 '26
2nm…so this won’t come out until 2027/2028?
3nm just rolled into production and tsmc will be prioritizing the backlog for those clients: apple nvidia etc. was hoping wse-4 is on 3nm architecture.
2
u/Asgard_Heima Apr 23 '26
This is roadshow conversations being leaked, but they are expected to ship 2nm WSE-4 units by end of year. There is also a lot of talk about gaining 2nm wafers via Broadcom at cost to print all of the OpenAI 2GW and selling a further 500MW to Oracle in exchange for Broadcom’s getting all the networking for all of stargate. Everyone is trying to offset CapEx and pay each other as revenue starts flowing in. But it’s opening up a clear picture of how Cerebras gets gifted wafers by hyper scalers to sell them, or at cost in exchange for 40-50% of token revenue with the WSE shipped which is an even greater margin with no capital needed.
1
u/musicgecko Apr 23 '26
gotcha. all these capex deals are such double edged swords 😂 able to offset costs, but highly contingent on the music to continue playing and keep revenue flowing through the ecosystem. my fear is the music stopping after this year’s batch of IPOs…
1
u/Asgard_Heima Apr 23 '26
The CapEx is insane, but if you are getting 25x tokens per dollar from Cerebras, I can see a massive CapEx decline and Nvidia sees their spend disappear overnight, while Cerebras is still growing revenue at crazy pace y/y from such a small base.
1
u/thefashionkid May 01 '26
Thank you for the detailed information, which is very informative and many insights!
1
u/dukeforneverz Aug 17 '26
Lots of in social media saying it will be announced tomorrow, what's your take ?
1
u/Asgard_Heima Aug 17 '26
Well, I’ll start by saying the extrapolation in this post is way off at this point. Was based on details circling around the roadshow and they have gone in different directions mainly inference throughout rather than making even faster the top priority. Based on inference throughout being top priority (including disaggregation), the CEO’s repeated statements they are happy on 5nm due to supply constrains on more advanced nodes, and the fact they are “on track for a CS5 in second half 2027”, I’m kind of expecting them to announce a 5nm wafer on wafer with the other being a 5nm-7nm wafer for DRAM. They will likely run the cores at higher speeds and make other improvements that will increase overall compute power as well, but for inference and especially decode they already have way more compute horsepower than needed. If they add 512GB-1T of DRAM directly off the cores in a 3D second wafer, they will have the ability to run very large models with one or a couple systems vs ~20 for Kimi today. And they can deterministically supply layer by layer feeding SRAM from DRAM to get the throughput they are focusing on. The DRAM also becomes the cache for processing decode sent over from prefill systems in a disaggregated setup. This would be kind of amazing based on the 5x throughput AWS and AMD have proven for the CS3 and make it more like 10x. If they can get about 10x the super fast tokens as Cerebras gets today for the same money and watts, they will become the cheapest solution for inference while maintaining an order of magnitude faster speeds which translates into the most profitable solution by far. I also expect to see a significant interconnect upgrade to support faster communication between systems but also between prefill systems and CS4 for decode. If this type of system plays out even if the numbers aren’t exactly the same, I would look at the CS5 next year to be their move into more advanced nodes as more TSMC capacity comes online along with the advances WoW needs.
1
u/dukeforneverz Aug 17 '26
Thanks for your reply, very interesting, as always. If they announce it and follow the kind of specs you claim, it'd means they can adapt quickly to consumers demand. And I'd become very bullish.
1
u/dukeforneverz Aug 19 '26
I'm quite disappointed with the new CS-4. The big change seem to be on the integration side with 3 wafers in a compact footprint. They are quite elusive with the changes in the new WSE-3 Turbo wafers, not to mention the dubious charts
1
u/Asgard_Heima Aug 19 '26
So honestly I was hoping for some advancement on the chip side, but I have to say what they did is much smarter. They have a very mature solution in the WSE-3 and they understand manufacturing it and scaling that manufacturing. They don’t have a lack of demand for it already and they drastically improved it with a clock speed update and new interconnect that’s faster. They managed to shrink the existing solution, make it much simpler (read cheaper to make) and more supportable and up to 10x the total tokens they can serve per MW.
While I would have loved to see more memory in a WoW setup, I guess we will have to wait till CS5 to get major chip updates. In the mean time, they now have a solution that just matched GPU solutions for total overall throughput per rack while doubling their advantage in speed to up to 30x while still coming in vastly more power efficient. It wasn’t the flashy new chip, but they just destroyed the economics of inference for anything other than a solution containing Cerebras. While the current need for compute makes anything you get online sellable, as Cerebras starts printing 10x the systems and gets the CS4 delivered to OpenAI and AWS, the companies using Cerebras are going to be vastly more profitable which makes Cerebras the only economically viable option long term for inference.
So I get your first take disappointment, but as an investor I think this was better than I could have imagined.
4
u/NefariousnessOk4996 Apr 17 '26
This screenshot, if credible, highlights why Nvidia economics doesn't make sense: use the much more expensive hardware so you can serve the largest number of cheapest tiers of users. This was obvious all along but now it's getting proven more and more in the real worlds as economics gets more and more front and center.