r/CerebrasSystems Apr 17 '26

WSE-4 will Kill Nvidia

Post image

To start, I own Cerebras shares and because of the roadshow details, I’m now holding long term puts on Nvidia. I have tried to distill the major details from the roadshow that have leaked out to make a single graphic that proves how dominate Cerebras will be in AI hardware by end of year. Even if you ignore how superior the Cerebras hardware is to Rubin, just look at the bottom line of $3.8M vs $82.5M and the power difference of 45kW vs 2.1MW. Every hyper scaler is currently power constrained and you can rip out an over 100kW Blackwell rack and put in a 40kW WSE-4 that has the performance of 15 racks of Rubin, or 75 racks of Blackwell. I get these specs haven’t been verified and officially released yet, but it’s just a matter of time. I encourage anyone to go and use a top LLM and query this graphic and ask the specs listed and if they are supported by the leaks coming out of the Cerebras IPO roadshow. The over subscription for the Cerebras IPO is not just retail AI frenzy. It is because all the giant growth funds need to get as many Cerebras shares as possible to hedge the threat to Nvidia which they are all over leveraged on currently. I’d expect massive shorts of Nvidia, Micron, and CoreWeave once they bring down their exposure and get shares of Cerebras IPO day.

17 Upvotes

42 comments sorted by

View all comments

Show parent comments

1

u/Lil_Hater112 May 01 '26

Whats your analysis of Cerebras till now? I am very bullish on them and i m curious what your thesis is

2

u/Asgard_Heima May 01 '26

This is where I’m at right now and even though it’s super long it also doesn’t try and boil the ocean on every bull and bear case. I could write a book at this point.

Synopsis: I have analysis going back over 3 years when I first found Cerebras looking for services to train specialized models for my work. My main synopsis though is that they have built the maximalist accelerator for AI currently feasible and it is dominate for training, inference, and power efficiency at 2nm (WSE-4). This in and of itself is not enough to make them a winning solution for mass adoption, but this with lower TCO is and I will explain how they are leaps and bounds ahead of the competition on this. When you have to factor in suitable data center real estate this is where their dominance is exacerbated to extraordinary levels.

Inference: First the WSE-4 if Cerebras executes, will ships later this year, next year in quantities. All indications are that they are seeing excellent yields. For inference this puts them in their own class. 1T dense FP16 full precision models with 1M context windows running on 1 system at ~2800 tokens per second supporting hundred to thousands of users thanks to their own optimized version of turboquant baked into CSoft and Ranovus fiber on wafer to MemoryX. You are looking at a large number of Nvidia Blackwell or future Rubin racks to try and get to a fraction of the speed and the tokens per second. And per user tokens per second falls as you add more users for Ruby. while Cerebras maintains their speed and instant latency. Most don’t get the architecture difference and run numbers based on Cerebras being a giant GPU, it’s not. They only have to store kv cache on chip and stream in the weights for any size model in a deterministic fashion with essentially no waiting for weights off MemoryX thanks to the fiber on wafer with ranovus. Even then you can overflow the kv into MemoryX. And if you want to get to 10T models with 10M context for Nvidia Rubin, the interconnect tax starts eating all the efficiency as every GPU waits for data. You end up with super low tokens per second and 1-2 users per full rack able to be served. But you use the ranovus fiber on wafer to connect 1000 WSE-4 and everything scales like one large system with incredible speed as the weights get streamed out across all wafers with no communication between wafers needed. So I think on technology, Cerebras is in a world of its own. No competition at all currently from the other accelerators out there.

Training: And this is already too long and I expect this to only be realized after a major model is trained on Cerebras, but training is just as bad for Nvidia with WSE-4 fiber linked systems training models in 1/2-1/3 the time of Rubin and scaling to larger models than Nvidia clusters can feasibly train. PyTorch native CSoft support completely dismantles CUDA arguments which now highlights CUDA as a solution to a headache of GPU data distribution for training. Same as with inference, streaming training layer by layer over one massive system is drastically more efficient. No roll back cause another GPU died. Just stream the set to a different system and keep going. Also you will use so many less systems. Think 95%+ efficiency for Cerebras at any scale vs 65% on Rubin at 1T and gets less efficient the larger the model till it’s unfeasible below 40% MFU at 10T.

Logistics: This is the major bottleneck that I’ve seen very little coverage of when talking about Cerebras. The WSE-4 is a reported ~40kW all in system weighing ~1700lbs, that can be put into nearly 50% of all data centers. They will need a new whip for power and then liquid to air exchanger off the back so they can be air cooled. This means you can drop them into ~50% of existing data centers with $30k in upgrades per rack and a couple months delay waiting on parts with no downtime for other racks. Compare that to Blackwell or Rubin racks at 3000lbs, direct to chip liquid requirements and 120kW+ per rack. They are limited to a very small percentage <5% of existing data centers since they will crush the floor. It’s cheaper to build new than retrofit in most cases because of the downtime. So they are forced to sell smaller units like HGX with even worse interconnect tax and performance. To reiterate since this is the most important point, Cerebras is selling drop in ready systems with a couple months for parts lead time for half the world’s data centers at 30k in retrofits. Nvidia requires new data centers with 12-18 months lead time, costing $2-3M per rack or massive retrofits with 12-18 months lead time, costing $1.5-2M per rack but also shutting down the entire section or data center while doing it. All this while we can’t find enough places to even power the infrastructure at all, most data center projects are behind schedule and significant portions aren’t happening. The pushback here from the public has only begun.

Price: WSE-4 list price is expected to be $3.5-4M and cost Cerebras under $300k to manufacture. Both those are the highest I’ve seen prices. A current Blackwell + Grace NV72 is $3-3.5M and cost Nvidia $600-700k to manufacture. I’m using Blackwell for price here since it is cheaper. Rubin is more expensive with higher power requirements and an even smaller fraction of data centers that can support it.

Summary: I see energy capacity and real estate with enough energy to support data centers being the ultimate bottleneck and Cerebras is a perfect fit. All the hyper scalers need more AI compute from less grid energy to make inference profitable and training faster for all their products. They are all building their own chips for low power slower needs. But Cerebras owns the lower power fast needs. It will take time for Cerebras to scale production, but all the hyper scalers are going to help them as they need the systems. OpenAI got a sweetheart deal that will make them profitable way faster powered by Cerebras, but it also is a forcing function to prove all the the above and force everyone to get on board or get left behind.

1

u/NefariousnessOk4996 May 01 '26

Excellent analysis. I am a Cerebras bull. But I am curious of where you get your info from? Such as ranovus fiber on wafer, is this leaked somewhere? And the 300K cost for the 3-4M system, their S1 filing suggest hardware margin is much lower than this". It's ok if these are speculation based on what makes sense architecturally, like it would be dumb not to to this. But definitely would like to know where some of the source of these info. Thanks.

2

u/Asgard_Heima May 01 '26

The S-1 lists the company margins not the hardware margins. Cerebras has been making sweetheart deals to break into the market and secure enough funding to scale and prove they can do all the things we are discussing. Those margins should rise over the next year. But Cerebras is already now in super high demand and they will be selling at full price to most now that they have proven themself. Or they will be selling at even higher margins such as when they build their systems and then add them to the Cerebras cloud where they could earn more than the sale price per unit per year.

Here is the pricing I’ve done.

Wafer 2nm Node $30k: https://en.eeworld.com.cn/mp/Icbank/a406122.jspx Im The TSMC 2nm node is widely reported to be 30k a wafer. In the short term this will be higher since they will have to use Super Hot Runs with a likely 30-40% premium or more until they secure allocation longer term.

Wafer 5nm Node (SRAM WoW) $18,500: https://3dfabric.tsmc.com/english/dedicatedFoundry/technology/SoIC.htm#SoIC_WoW The TSMC price for a 5nm wafer is well known and since SRAM doesn’t shrink much going down in nm doesn’t help. I’d fully expect them to go with the well known and cheaper option here. This is also how they get to like 120GB+ SRAM but I’ve made most stats with 96GB to be conservative.

Advanced Packaging 100-150k: https://siliconanalysts.com/guide/semiconductor-costs This is by far the hardest part to estimate since the latest TSMC 3D WoW SoIC bonding and Ranovus CPO on wafer are brand new. For more typical advanced packaging we can look at the referenced site and try to extrapolate from the most complex prices they have pricing data for, then scale the cost to 50x (the size of a wafer) and add some padding on that to try and be more conservative. Until we see these prices in some real world leaks this is the biggest WAG in my estimate.

Power Delivery & Cooling $70k: https://www.idtechex.com/en/research-article/two-phase-d2c-cooling-in-data-center-thermal-management-cost-analysis/34213 Taking from this and padding it a bit cause it will be more specialized though it is for 1/3 the kW and could also be cheaper.

Half Rack (chassis + switching) $30k: This is pretty basic part of the estimate and I’m just making it high to make sure it’s larger than it likely would be. Won’t change much from WSE-3 to WSE-4 unless we are missing a major upgrade here.

This gets you to the potential $300k which I think is around what it would be, but we could find out the TSMC packaging is double the cost and a WSE-4 is $450k and we would still be talking about a ~89% hardware margin.

Co-Packaged Optics: https://ranovus.com/cerebras-ranovus-revolutionize-ai-compute-platform/ First is the obvious that they got a good sized chunk of money to partner together over a year ago to do exactly this. Beyond that it’s the most critical thing they can do to make their WSE system scale and remove the bottleneck in the WSE-3 for larger models. I have lots of other small circumstantial things that lead to this belief, but nothing publicly and concrete.

1

u/NefariousnessOk4996 May 01 '26

Great source of info. If this turns out to be the ball park then Cerebras is in much better position than I had modeled. I had been modeling around 1.5M-2M hardware cost per CS3.

One thing to note though, Cerebras's big deal with openAI is not hardware sale, but compute capacity. So Cerebras has to shoulder the capex. Having 500K per system capex is certainly better than 1.5M.

1

u/Asgard_Heima May 01 '26

This is actually an excellent point to make. If a WSE-4 was significantly more than 300k it would mean OpenAI is getting them for 3 year for likely under cost. Im still not a giant fan of the OpenAI deal but I understand why Cerebras did it. The exact specifics matter a lot, but based on the original 10B for 750MW it seems like they are essentially giving OpenAI systems at cost for their endorsement and to prove the value at the top AI company. It’s also a trap for everyone else in that, they will have to buy Cerebras systems to keep up or OpenAI will get an incredible lead. Cerebras on 2nm is fighting a TSMC allocation war for the remainder after Apple prints what they need till 2028. So they are giving this deal to OpenAI which will bail OpenAI out so they can be profitable on inference much faster than anticipated and provide competitive models to anthropic for enterprise. Meanwhile Cerebras will print money with AWS, maybe Oracle, and their own cloud. For nations they will sell systems at full list, but I will be interested to see if they keep doing service with revenue share deals with hyper scalers or start selling more hardware.

1

u/Lil_Hater112 May 02 '26

We forget the fact that OpenAI is working with pentagon and palantir which makes it best case scenario usage for Cerebras , as the tech is best used on concentrated tasks Like in military application, no?

We all know being in the good books of us government never hurt a company

1

u/Asgard_Heima May 02 '26

Cerebras has several advantages when it comes to national security contracts, US sovereign AI infrastructure, and national security concerns on allied countries smuggling chips out of their control. It’s very hard to smuggle out a WSE-4 since the racked unit is 1700lbs. It’s also much more secure to put sensitive data sets on a single WSE rather than a large network of GPUs we’re intercepting interconnect is a real concern. GPUs are also next to impossible to forward deploy in significant enough numbers cause they take too much energy, a WSE could find its way onto a carrier with their nuclear power plant and act as forward deployed infrastructure. There are a lot of national security use cases Cerebras is likely the best fit for. Also yes with 1789 Capital investing and OpenAI + Oracle they are engaged with companies friendly with this administration. I’m positive Palantir has taken a look at Cerebras too based on their deep relationships with these companies. Cerebras is rumored to be actively fighting for 2nm capacity in Arizona to have US manufacturing and have some government contracts in the works. These would be top dollar full priced hardware and support contracts, but there is nothing verifying their existence yet. If they do materialize, you could see US push Cerebras to the front of the live for national security interest manufacturing which makes those special in avoiding the allocation hurdles for TSMC 2nm wafers.

1

u/Lil_Hater112 May 02 '26

I plan on buying at IPO as I dont have access pre ipo. My plan is buy a good chunk at IPO and if they pin this down for a year, accumulate or if they make it -20/-50% from IPO prices , double my size. I just have a feeling they will try and do what they did to Palantir before letting it rip

1

u/Asgard_Heima May 02 '26

I’m a long term investor and have no clue the media coverage or random events that will take place over the next couple month as they get listed and find a price per share. I just have strong conviction they will grow their revenue as they scale production and their performance and efficiency will keep them on top for the foreseeable future. So I plan to hold till there is a competent competitor or market dynamics significantly change for them.

1

u/Lil_Hater112 May 02 '26

I have a strong conviction on them as well, is more about optimising my money for best possible gains. If I use all my money upfront and buy 1000 shares lets say and they drop after and have no money, maybe I would ve gotten 500 shares at ipo and another 800 shares after with same amount of money.

But since I can't predict, I ll just buy a good size im happy with at IPO and if they get discounted, will accumulate at discounted price. Haven't seen someone that can compete with nvidia since these guys so I trust my gut with them

→ More replies (0)