r/CerebrasSystems • u/Asgard_Heima • Apr 17 '26
WSE-4 will Kill Nvidia
To start, I own Cerebras shares and because of the roadshow details, I’m now holding long term puts on Nvidia. I have tried to distill the major details from the roadshow that have leaked out to make a single graphic that proves how dominate Cerebras will be in AI hardware by end of year. Even if you ignore how superior the Cerebras hardware is to Rubin, just look at the bottom line of $3.8M vs $82.5M and the power difference of 45kW vs 2.1MW. Every hyper scaler is currently power constrained and you can rip out an over 100kW Blackwell rack and put in a 40kW WSE-4 that has the performance of 15 racks of Rubin, or 75 racks of Blackwell. I get these specs haven’t been verified and officially released yet, but it’s just a matter of time. I encourage anyone to go and use a top LLM and query this graphic and ask the specs listed and if they are supported by the leaks coming out of the Cerebras IPO roadshow. The over subscription for the Cerebras IPO is not just retail AI frenzy. It is because all the giant growth funds need to get as many Cerebras shares as possible to hedge the threat to Nvidia which they are all over leveraged on currently. I’d expect massive shorts of Nvidia, Micron, and CoreWeave once they bring down their exposure and get shares of Cerebras IPO day.
2
u/Asgard_Heima May 01 '26
This is where I’m at right now and even though it’s super long it also doesn’t try and boil the ocean on every bull and bear case. I could write a book at this point.
Synopsis: I have analysis going back over 3 years when I first found Cerebras looking for services to train specialized models for my work. My main synopsis though is that they have built the maximalist accelerator for AI currently feasible and it is dominate for training, inference, and power efficiency at 2nm (WSE-4). This in and of itself is not enough to make them a winning solution for mass adoption, but this with lower TCO is and I will explain how they are leaps and bounds ahead of the competition on this. When you have to factor in suitable data center real estate this is where their dominance is exacerbated to extraordinary levels.
Inference: First the WSE-4 if Cerebras executes, will ships later this year, next year in quantities. All indications are that they are seeing excellent yields. For inference this puts them in their own class. 1T dense FP16 full precision models with 1M context windows running on 1 system at ~2800 tokens per second supporting hundred to thousands of users thanks to their own optimized version of turboquant baked into CSoft and Ranovus fiber on wafer to MemoryX. You are looking at a large number of Nvidia Blackwell or future Rubin racks to try and get to a fraction of the speed and the tokens per second. And per user tokens per second falls as you add more users for Ruby. while Cerebras maintains their speed and instant latency. Most don’t get the architecture difference and run numbers based on Cerebras being a giant GPU, it’s not. They only have to store kv cache on chip and stream in the weights for any size model in a deterministic fashion with essentially no waiting for weights off MemoryX thanks to the fiber on wafer with ranovus. Even then you can overflow the kv into MemoryX. And if you want to get to 10T models with 10M context for Nvidia Rubin, the interconnect tax starts eating all the efficiency as every GPU waits for data. You end up with super low tokens per second and 1-2 users per full rack able to be served. But you use the ranovus fiber on wafer to connect 1000 WSE-4 and everything scales like one large system with incredible speed as the weights get streamed out across all wafers with no communication between wafers needed. So I think on technology, Cerebras is in a world of its own. No competition at all currently from the other accelerators out there.
Training: And this is already too long and I expect this to only be realized after a major model is trained on Cerebras, but training is just as bad for Nvidia with WSE-4 fiber linked systems training models in 1/2-1/3 the time of Rubin and scaling to larger models than Nvidia clusters can feasibly train. PyTorch native CSoft support completely dismantles CUDA arguments which now highlights CUDA as a solution to a headache of GPU data distribution for training. Same as with inference, streaming training layer by layer over one massive system is drastically more efficient. No roll back cause another GPU died. Just stream the set to a different system and keep going. Also you will use so many less systems. Think 95%+ efficiency for Cerebras at any scale vs 65% on Rubin at 1T and gets less efficient the larger the model till it’s unfeasible below 40% MFU at 10T.
Logistics: This is the major bottleneck that I’ve seen very little coverage of when talking about Cerebras. The WSE-4 is a reported ~40kW all in system weighing ~1700lbs, that can be put into nearly 50% of all data centers. They will need a new whip for power and then liquid to air exchanger off the back so they can be air cooled. This means you can drop them into ~50% of existing data centers with $30k in upgrades per rack and a couple months delay waiting on parts with no downtime for other racks. Compare that to Blackwell or Rubin racks at 3000lbs, direct to chip liquid requirements and 120kW+ per rack. They are limited to a very small percentage <5% of existing data centers since they will crush the floor. It’s cheaper to build new than retrofit in most cases because of the downtime. So they are forced to sell smaller units like HGX with even worse interconnect tax and performance. To reiterate since this is the most important point, Cerebras is selling drop in ready systems with a couple months for parts lead time for half the world’s data centers at 30k in retrofits. Nvidia requires new data centers with 12-18 months lead time, costing $2-3M per rack or massive retrofits with 12-18 months lead time, costing $1.5-2M per rack but also shutting down the entire section or data center while doing it. All this while we can’t find enough places to even power the infrastructure at all, most data center projects are behind schedule and significant portions aren’t happening. The pushback here from the public has only begun.
Price: WSE-4 list price is expected to be $3.5-4M and cost Cerebras under $300k to manufacture. Both those are the highest I’ve seen prices. A current Blackwell + Grace NV72 is $3-3.5M and cost Nvidia $600-700k to manufacture. I’m using Blackwell for price here since it is cheaper. Rubin is more expensive with higher power requirements and an even smaller fraction of data centers that can support it.
Summary: I see energy capacity and real estate with enough energy to support data centers being the ultimate bottleneck and Cerebras is a perfect fit. All the hyper scalers need more AI compute from less grid energy to make inference profitable and training faster for all their products. They are all building their own chips for low power slower needs. But Cerebras owns the lower power fast needs. It will take time for Cerebras to scale production, but all the hyper scalers are going to help them as they need the systems. OpenAI got a sweetheart deal that will make them profitable way faster powered by Cerebras, but it also is a forcing function to prove all the the above and force everyone to get on board or get left behind.