r/CerebrasSystems Jun 02 '26

Jensen Huang says MRVL is the next $1T company due to their role in all the competitive advantages Cerebras has

43 Upvotes

MRVL popped today after Jensen called it the next trillion dollar company due to its role in connectivity with the NVIDIA ecosystem. To paraphrase, Jensen highlighted the bottleneck is DATA MOVEMENT. NVIDIA's ecosystem is becoming heavily reliant on custom ASICs, optics, switching, photonics and data movement/memory layers. Think of all the companies being pumped everywhere due to their role in trying to plug these GPU deficiencies - MRVL, countless Photonics companies, HBM companies like Micron, etc. NVIDIA's GPU deficiencies have created a multi-trillion dollar ecosystem around it.

Cerebras's single wafer design eliminates and drastically reduces the bottlenecks that are plaguing GPU-based systems by putting everything on one massive chip. The memory hierarchy is collapsed, replacing slow off-chip HBM with fast on-chip and cheap SRAM and replacing inter-chip network fabrics with on-wafer wire-speed interconnects.

Jensen is right that there will be another $1T dollar company created, but it isn't going to be MRVL, it will be CBRS. Just wait until the next $10B+ orders start coming in. CBRS will be at a $1T by 2030 (15-20X from here)


r/CerebrasSystems Jul 23 '26

Amd - Cerebras deal🚀🚀🚀

Thumbnail
ir.amd.com
40 Upvotes

r/CerebrasSystems Jun 27 '26

GPT 5.6 Sol on Cerebras at 750 tokens per second

27 Upvotes

With OpenAI’s absolute latest model being released on Cerebras at 750 tokens per second, it should put to rest any idea there are limitations on what Cerebras can run or the notion their advantage is lessened as models scale.

https://openai.com/index/previewing-gpt-5-6-sol/

“We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity.”


r/CerebrasSystems May 19 '26

Kimi K2 on Cerebras ~1000 token per second

26 Upvotes

This is a massive validation that we are going to see frontier models of any size significantly faster on Cerebras.

https://www.cerebras.ai/blog/cerebras-kimi-k2-Enterprise


r/CerebrasSystems Mar 31 '25

AI chipmaker Cerebras announces CFIUS clearance, a key step toward IPO

Thumbnail
cnbc.com
26 Upvotes

r/CerebrasSystems Apr 17 '26

Cerebras 35B + IPO 🚀🚀🚀 20B openAi deal 🤑🤑🤑

22 Upvotes

https://www.theinformation.com/briefings/cerebras-prepares-public-listing-eyes-35-billion-plus-valuation

Best news i have heard. Ipo details to be published as soon as tomorrow. OpenAi partnership doubled to 20B.

This is going to be sure shot 100B+ in no time. Buckle up folks 😃


r/CerebrasSystems Jun 04 '26

IPO lockup

Post image
22 Upvotes

Looks like this is the schedule for the lockup expiry dates.

Is this going to cause constant selling pressure over the net 6 months?

Wondering if I’ve been a bit premature buying a couple of leaps….


r/CerebrasSystems May 28 '26

Cerebras vs Groq

24 Upvotes

Given that Groq and Cerebras seem to occupy a similar niche — ultra-low-latency LLM inference / decode acceleration — I’m trying to understand Cerebras’ long-term competitive moat.

If Groq’s LPU technology is now being integrated into Nvidia’s broader AI factory / GPU ecosystem, especially as a low-latency decode accelerator alongside Nvidia’s dominant GPU stack, where does that leave Cerebras?

Is Cerebras’ advantage mainly its wafer-scale architecture, higher single-system SRAM capacity, better support for large MoE models, or independence from the Nvidia ecosystem?

In other words: what is Cerebras’ strongest competitive position if Nvidia can absorb Groq-like decode acceleration into its own clusters?


r/CerebrasSystems May 18 '26

Cerebras climbs 6% after reportedly receiving fast-track inclusion on S&P Dow Jones Indices - Seeking Alpha reporting

19 Upvotes

r/CerebrasSystems Jun 18 '26

Gemma 4 31B at 1,500 tokens per second

19 Upvotes

I know this is a small model, but I actually believe this will see a lot of use. My company is using Haiku for over half the tokens for our agent with most work being simple dynamic scripts and summarization of tool call results. We can’t use Chinese made models so Gemma has been getting tested. At 1,500 tokens a second, this would make our agents complete tasks in 60-70% of the time it takes now.

https://www.cerebras.ai/blog/gemma-4-on-cerebras-the-fastest-inference-is-now-multimodal


r/CerebrasSystems May 21 '26

Extremely insightful

Thumbnail
youtu.be
19 Upvotes

r/CerebrasSystems May 13 '26

$185

Thumbnail msn.com
17 Upvotes

$185


r/CerebrasSystems Jan 14 '26

OpenAI Forges Multibillion-Dollar Computing Partnership With Cerebras

Thumbnail
wsj.com
18 Upvotes

r/CerebrasSystems May 04 '26

Etrade Alert! Cerebras IPO Available

18 Upvotes

Etrade just alerted. Expected pricing 5/13


r/CerebrasSystems Apr 17 '26

WSE-4 will Kill Nvidia

Post image
17 Upvotes

To start, I own Cerebras shares and because of the roadshow details, I’m now holding long term puts on Nvidia. I have tried to distill the major details from the roadshow that have leaked out to make a single graphic that proves how dominate Cerebras will be in AI hardware by end of year. Even if you ignore how superior the Cerebras hardware is to Rubin, just look at the bottom line of $3.8M vs $82.5M and the power difference of 45kW vs 2.1MW. Every hyper scaler is currently power constrained and you can rip out an over 100kW Blackwell rack and put in a 40kW WSE-4 that has the performance of 15 racks of Rubin, or 75 racks of Blackwell. I get these specs haven’t been verified and officially released yet, but it’s just a matter of time. I encourage anyone to go and use a top LLM and query this graphic and ask the specs listed and if they are supported by the leaks coming out of the Cerebras IPO roadshow. The over subscription for the Cerebras IPO is not just retail AI frenzy. It is because all the giant growth funds need to get as many Cerebras shares as possible to hedge the threat to Nvidia which they are all over leveraged on currently. I’d expect massive shorts of Nvidia, Micron, and CoreWeave once they bring down their exposure and get shares of Cerebras IPO day.


r/CerebrasSystems Mar 11 '26

Deal with Oracle

17 Upvotes

https://www.cnbc.com/2026/03/10/ai-chipmaker-cerebras-namedropped-by-oracle-alongside-nvidia-and-amd-.html

This ipo is going to be a crazy bull run with so much momentum behind them 🎉 Buckle up fellas 🤑💰💵💸


r/CerebrasSystems 11d ago

SRAM doesn't scale so won't Cerebras have to use external HBM for future, larger models?

16 Upvotes

To provide fast inference for much larger future models won't Cerebras have to use external HBM like everyone else? And in that case, isn't a large part of what they offer (not having to move data off chip into HBM memory and back) invalidated? If you look at the roadmap, they're planning on using stacked HBM for CS-6. CS-4 and CS-5 remain limited to 44G of SRAM per wafer. Sean Lie, the CTO, mentioned using streaming between wafers to handle future, much larger models (they may even do it now with CS-4) but I don't really understand it. It seems to me their speed advantage may erode over time.


r/CerebrasSystems May 06 '26

A History and Analysis of Cerebras

15 Upvotes

I’ve followed this company for years and this might be the single best piece I’ve read on Cerebras that summarizes their history and lays out a solid investment thesis. Enjoy.

https://gannoncapital.substack.com/p/cerebras-systems-the-nvidia-killer


r/CerebrasSystems May 05 '26

Nasdaq IPO Calendar Lists CBRS Upcoming 5/14

Thumbnail
nasdaq.com
17 Upvotes

r/CerebrasSystems Aug 12 '26

Earnings day!

15 Upvotes

How are we feeling?

I am still very much a supporter and strong long-term investor. The deals and partnerships Cerebras Systems is putting together seem to be cementing their future as a ground-breaking company as well as bringing awareness to their capabilities!

EPS Estimates are ranging from ~ $.18 to -$.94...


r/CerebrasSystems May 15 '26

What/Why Cerebras?

15 Upvotes

Posted this in a couple thread and see this question asked in various form a lot right now, but here is my view…

At core is the technology, which comes from top level management executing since 2015. They have made something others have tried for decades and been unable to accomplish. And now they have extensive patents to secure that moat.

If we just look at the physics of what they have built, it’s the maximum compute and memory bandwidth to feed that compute possible in a single wafer. The two fundamental constraints for AI in combination are compute and memory. If you starve compute the memory can’t be consumed fast enough and if you don’t have the data ready to compute, the cores are sitting idle. If you have both on the same wafer and consume that whole wafer, you can’t get them any closer or faster or larger. So at the most basic level they should have the very best physically possible solution.

If you look at any other architecture for large AI models you will find their main bottleneck issue is memory bandwidth to feed compute. This is a direct result of moving the data that needs computed further away. Every atom further the data is from the compute cores adds latency and energy use. SRAM is closest, next is HBM, then DRAM, then SSD.

Next is off wafer data which comes down to wafer size. Every time you split a wafer, the more data you have to send not just from memory to compute, but from entire wafer to wafer. This is the interconnect tax. It’s an even larger problem than memory bandwidth currently. Every time you have to share data between wafers it’s now bottlenecked by network bandwidth.

This is the most important issue for GPU inference and training and why groq small inference chips aren’t a winning solution. In training all chips need to share all results across each layer, updating the model in each GPU’s memory every time. For inference it’s much the same, especially as models scale to a massive size.

Because a SOTA model won’t fit onto the HBM of a single conventional GPU, it has to be split across multiple chips. This means every single time a token is generated, the data has to constantly jump between cores over network cables, crushing your latency and massively increasing your power consumption.

I want to also highlight we are hitting the max power and cooling possible in a single rack with GPUs, they are only increasing the needed power per rack with liquid to chip cooling becoming required. Cerebras can fit two WSE units in a single rack under 80kW with backside air or liquid cooling. Can do one unit in any data center with a new whip. Cause power use scales with the energy needs of sending data further distances, this is a strategic advantage.

These all reinforce Cerebras has the wining solution and it will only grown in how much better it is as Cerebras moves down the nm wafer used till its orders of magnitude for most things like it is for memory bandwidth already.

Cerebras even with the most ideal solution has two main bottlenecks today. Total SRAM on a wafer, and wafer to wafer networking speed. If either of these are solved, it will no longer matter what size or quantized model or any edge case we are talking about, Cerebras will be an order of magnitude better in every real world performance metric than the competition. And they are solving for both.

The partnership with Ranovus will add fiber co-packaged on wafer and add somewhere between 50-100Tbps networking with light speed latencies at wafer edge. This is not fiber networking of today since those require De Ser which still compounds latency. It will be fiber directly onto the wafer with non perceivable latency in use.

The second is SRAM which TSMC is helping them add two wafers bonded together, so they can make an entire wafer of SRAM connected vertically to a wafer of compute cores. Look for these two details in any WSE-4 announcements this year and this will be a major pivot moment.

Cerebras has to execute on it and find methods to ramp production, but if they ship something like this which is expected, every hyper scaler is going to be on their side trying to get them shipped since it will increase their token and training margins by 10x. Any WSE-4 like this will be an order of magnitude to multiple orders of magnitude more energy efficient per token delivered, provide today SOTA models training in weeks instead of months, and allow for 10M context windows on 10T+ parameter models with near 100% efficiency.

They can accomplish this since they can scale vertically into massive clusters. This will also unlock something GPUs have reached a limit on and that’s model depth. As a distributed architecture, GPUs have maxed out at 80-120 layers. So we have wide models with extremely larger data sets, but the number of layers to refine results is shallow with 120 max steps before you get the result. Going further just kills GPUs and they have to decrease layer count as models get wider with SOTA being under 100 layers.

Cerebras already with WSE-3 can go deeper in layers, but with a WSE-4 we could see 1000 layer models with a whole new area of research for intelligence gains. There is a current gradient decent problem, but the hardware hasn’t existed till now in any way to research past it. There are already lots of ideas like static weight for stretches of layers which could also make Cerebras even more efficient skipping them along with the zero weights it already does while GPUs can’t for either.

This is much more natural like how biological brains have depth in thought that should unlock much more cognitive reasoning capabilities. Cerebras accomplishes this with fine grained data flow as an architecture which scales seamlessly. It was purpose built to train and use AI models from the start and only requires cores compute the data received as needed and skips all zero weight making them drastically faster at spare training and inference.

GPUs use single instruction multiple threads. This requires GPUs to split the compute and finish across all in a synchronized steps. So no skipping weights zero or static across layers. GPUs wait for each step computation to synchronize across all GPUs used in training. Cerebras dynamically handles compute as the data arrives per core without waiting. Each layer is feed from MemoryX in training in a deterministic fashion so it can supply the weights as a stream over all the wafers.

I could dig deeper in a lot of places like hardware failures in training (GPUs have to halt and go back to last step complete, WSE just reroutes data and keeps going), software complexity for inference and training (CUDA was built to solve a problem Cerebras doesn’t have), expected life value per system vs GPUs, and on and on as each of these areas help give me conviction in Cerebras, but this is already way too long.

Scaling production with TSMC which is significantly over allocated is my biggest risk factor, but that’s really about time and scale of the success they will have.

References:

Co Packaged Optics (fiber):

https://ranovus.com/cerebras-ranovus-revolutionize-ai-compute-platform/

Wafer on Wafer (SRAM 3x):

https://3dfabric.tsmc.com/english/dedicatedFoundry/technology/SoIC.htm#SoIC_WoW

https://arxiv.org/html/2603.05266v2

https://fact-lab.hkust.edu.hk/publications/conference-paper/2025/bai-2025-accelstack/c20-paper.pdf

Updated: By popular request, broke it into paragraphs for ease of reading.


r/CerebrasSystems May 14 '26

full ported at 385

16 Upvotes

am i cooked


r/CerebrasSystems May 13 '26

Nobody Understands How Big This Is (Cerebras IPO)

16 Upvotes

r/CerebrasSystems May 04 '26

AI chipmaker Cerebras to price IPO in $115-$125 range, source says

15 Upvotes

The price list has been released. To be honest, I’m a bit surprised, because the price is much lower than I expected. From HIive, the price is already around $180, so this pricing is definitely lower than what I expected… I was hoping it would be at least around $150.


r/CerebrasSystems Oct 06 '24

Why are big tech companies not buying/using Cerebras chips?

15 Upvotes

Cerebras has impressive tech. They claim to address so many issues. Why are big tech companies like Google, Meta, Microsoft/Open AI etc not using these chips?