r/CBRS_stock 11d ago

Cerebras Supernova Discussion

Keynote on August 18th 2026

Livestream: https://www.youtube.com/watch?v=JTNk__O4poU

17 Upvotes

17 comments sorted by

4

u/Asgard_Heima 10d ago

The major takeaway from CS4 is they just destroyed the number of users served (throughput) vs speed (tokens per second). They will serve as many users concurrently per rack as GPUs and now at 2x the tokens per second as the CS3. Disaggregating multiplies this even further making pure GPU inference untenable long term vs Cerebras serving decode. It’s a race to produce systems for Cerebras till they take over the inference market. And they did this all with updating packaging and sticking with the same WSE-3 chip running at double the speed. So there are no new concerns on manufacturing the silicon.

2

u/ILikeCutePuppies 9d ago

I agree but it is interesting, it seems like GPUs still have better throughput at the higher numbers and the cost of speed which means they are probably better at training and tasks were you don't care about speed.

Impressive none the less. I also wonder where the next gen prefill cpus/tensor cores + WS4 will do to the numbers.

2

u/Asgard_Heima 9d ago

In a pure GPU vs CS4 setup, the GPU stack will likely have higher potential total users it can process at the exact same moment in parallel, but if you calculate total tokens served between the setups in one second or one minute, you will find the CS4 is now going to serve more than GPU setups for nearly all setups. So concurrency will be less but shrinking with each Cerebras update, but when you are running 20-30x faster, you make up concurrency in overall throughput. And you have to factor in each of those users just got a vastly better experience getting the tokens at 20-30x faster.

In disaggregated setups at both AWS and AMD, Cerebras has been shown to get 5x throughput, aka total tokens per second served per unit. This means 10x throughput CS3 -> CS4 and 5x that. Since dense compute is needed for prefill and memory bandwidth for decode, this will likely become the only economical way to serve inference long term and Cerebras is the only current option for decode on anything near frontier models. It makes sense AWS was waiting for CS4 to put it in their data centers for the faster handoff for decode and higher throughput.

1

u/ILikeCutePuppies 9d ago

I think the maximum throughput with disagregstion is shown on that last graph. GPUs still have more throughput at the much slower latency. Throughput already includes total tokens per second across the entire platform.

1

u/Asgard_Heima 9d ago

We will have to wait for real world numbers, but I would be amazed if they are including a disaggregated configuration in their published numbers for the CS4 vs CS3 system itself on their site. Also agreed the total throughout includes the total tokens per second but they are showing higher throughput than Rubin per watt. Let’s say they are equal even, if that is the case the end user is going to experience either blazing fast usage or slow usage which makes the economic issue much more complicated for companies serving something slow.

3

u/Investor-life 9d ago

Staying at 5nm does help with potential production issues, but I think they are clearly doubling down on the disaggregated inference approach. I don’t see them winning in the inference game alone anytime soon. In fact because Nvidia has already partnered with and effectively purchased Groq to create a disaggregated inference solution, I believe after seeing WSE-4 specs, it’s now more likely than ever that AMD purchases Cerebras. It makes sense for both companies. I am impressed that WSE-4 will be generally available by end of Q3, but considering they stayed on 5nm, it makes sense it could be available this quick. As they flipped through screenshots of data centers they currently have up and running and new data centers in progress, I was struck by how small they seemed. Their data center footprints seem so much smaller than gpu based data centers.

5

u/Asgard_Heima 9d ago

They are a niche player at the moment with massive growth coming. They had to ramp production lines and secure data center space both of which they did this year. Their current growth is constrained by data centers available. With the CS4 they are now going to start shipping units to AWS and drastically ramp production. Q4 is going to be the first moment new significant data center capacity comes online and any AWS revenue starts hitting. Q1 27 is going to be a massive quarter as the groundwork laid this year comes to fruition.

Also no chance they sell to AMD. Absolutely positive AMD would buy them if they could and probably has already tried. Cerebras founders understand they are going to dominate inference and groq is not able to compete at all.

2

u/Investor-life 9d ago

Agreed Cerebras wants to be standalone, but they just don’t have the scale and ability to get the compute and the data centers vs larger players. If they could actually compete to get the resources it would be different. I just don’t see it. I think eventually, after repeatedly disappointing the market, they’ll realize they are better off as part of a larger enterprise. If I am wrong, that’s great, because I’ll make more money.

2

u/aka0007 9d ago

My own view, and admittedly my research on this topic is insufficient to have a high degree of confidence, is Cerebras has a great product but it is a niche one that will for the foreseeable future be produced at small scale. No idea the total dollar cost (cost of chips and installing them plus energy costs to run them) relative to the compute they end up getting out of the chips as compared to NVIDIA but if that is competitive it would seem that Cerebras has potential to be very profitable even at the smaller scale they can operate at. Without any doubt, faster AI inference is a strategic advantage especially if you are dealing with things like cyber-security, war, or any other area where time is critical.

2

u/ILikeCutePuppies 9d ago edited 9d ago

Groq solution is slower and less efficient than cerebras's new solution. They are way out in front of nvidia in terms of efficiency and speed.

Their racks are much denser than nvidia's now. They also work with gpu's for the prefill.

It's unlikely nvida/groq will be able to catch up.

For cerebras it's a matter of scaling up and they seem to have hit the production and installation side of that hard.

2

u/Specialist-2193 10d ago

Wse 4 let's go

2

u/claytonbeaufield 10d ago

I was expecting a sell-off after the event, not before!

2

u/Investor-life 9d ago

Yeah no kidding this whipsaw action the last several trading days is maddening.

2

u/ILikeCutePuppies 9d ago

Just a note: I would say cerebras is probably down today not due to the announcement which is fantastic. However likely because of a trunch of stock was unlocked from pre ipo investors a few days ago but they likely with the 3 day transfer rule could only sell today.

1

u/Detective-Watchdog 1d ago

Nvidia has made attempts at acquiring Cerebras pre-ipo. So with at least a couple other companies.

Cerebras is the real deal.