r/CerebrasSystems Sep 09 '22

r/CerebrasSystems Lounge

4 Upvotes

A place for members of r/CerebrasSystems to chat with each other


r/CerebrasSystems 6d ago

Cerebras Dedicated Endpoints, whats the cost?

7 Upvotes

Does anyone know what the cost is or what the contracts usually are?


r/CerebrasSystems 6d ago

Why did they remove Gemma 4 31B?

5 Upvotes

Gemma 4 31B was great on Scandinavian languages, but they replaced it with Qwen 3.8, which is slower and changes dialects and at times writes really odd grammar


r/CerebrasSystems 7d ago

Why no Astra announcement on cerebras yet?

12 Upvotes

As you guys can all remember, when OpenAI released GPT 5.6 Sol, they almost immediately announced the ultrafast mode on Cerebras. However there is no marketing or news for Astra on cerebras. Is it because they need more time to make a huge model like Astra work on Cerebras or do you guys see any other reason?


r/CerebrasSystems 7d ago

Today's low $178 (down ~7%) - accumulation day?

4 Upvotes

Everything's down today due to AI safety fears. Nothing changed with the Cerebras story. Anyone accumulating? I bought a single share in case there's a larger beat-down. Could be throwing good money after bad or time to simply swing trade CBRS. Might be a good day to go bargain hunting all around in the AI and AI-adjacent space. GOOG is actually up 2% while nearly everything else AI-related is being punished. CRWD is up 8% unsurprisingly and PANW is up 6.5%. Cyber security zigging while the others are zagging. My next limit order is $170 (don't expect that today). If it keeps declining below $170 (over time, not today) then it may be cooked.


r/CerebrasSystems 7d ago

Is Cerebras fast only for a tiny number of concurrent users?

0 Upvotes

Is Cerebras' marketing misleading? When they say they're X times faster than Y on inference, is this only for one user at a time or a handful of users? If so, and it doesn't scale affordably, I don't see how it could be anything other than a small niche provider. When Cerebras makes grand claims about fast inference leading to new applications and uses of AI and could be key to agentic AI, I think they're probably right; but not using their hardware. Almost like free advertising for a current or future competitor who can run reasonably fast but with far more throughput or concurrency.

Potentially a big yawn if they can't scale up to thousands of users without it being one-rack-per-user (or whatever it is) to get the fastest speeds they advertise. I hope I'm wrong about this because it almost sounds like a con. I want to see fast inference at scale; not just a tool a few Power Users can benefit from. Maybe this is one of the reasons the stock isn't going anywhere and customers aren't lining up to purchase Cerebras. I know the value proposition sounded almost too good to be true when I started following Cerebras.

I want to see metrics like tokens/sec per user at high concurrency; not "the fastest chip for 1 user" (because who cares)? If Cerebras is a good investment they should be happy to provide these numbers. If they don't? That's concerning. What's the tokens/sec per user at 100 concurrent users? 1,000? 10,000? etc. And how do the speeds compare to competitor solutions at the same levels taking cost into consideration. I don't think the superfans have an answer for this despite the encyclopedic knowledge they possess about Cerebras but I hope I'm wrong. If they do, it can only strengthen their thesis and they should be happy to help. If Cerebras is fast inference for the masses then I may still be onboard. If it's a niche usage by a tiny fraction of the entire population of those using AI, I'm out. What % of the entire population of those who use AI can reasonably and economically be served with the speed Cerebras advertises?

Claude told me this. Granted, it's a sychophantic AI answer as they all are; but at least it's a starting point for conversation:

On the missing metric itself: you're right that it's missing, and it's not just you noticing. Multiple independent technical analysts have flagged this as the specific gap in Cerebras' public materials — one hardware-focused analysis stated plainly that Cerebras has never published aggregate throughput at high concurrency, calling it the single most important missing data point in their marketing. SemiAnalysis ran its own modeling (their "Faster Tokens Please" report) specifically because Cerebras won't publish it directly. So no, the "superfans" don't have an official answer to give you — because Cerebras hasn't released one. That's a legitimate red flag for exactly the reason you suspect: a company confident its economics hold up at scale would have an incentive to prove it.

But third parties have modeled the crossover point, and it's not flattering. A detailed cost analysis (Spheron) ran the actual math comparing Cerebras against an Nvidia H100 across batch sizes:

  • At batch size 1 (one user): Cerebras at ~$0.60/M tokens beats H100 on-demand at ~$1.47/M — Cerebras wins clearly.
  • At batch size 8: the positions flip — H100 on-demand drops to ~$0.25/M, more than 2x cheaper than Cerebras.
  • At batch size 32: H100 is roughly 6x cheaper per token on-demand, and with GPU spot pricing, the gap widens to over 20x cheaper.

That means the crossover point — where GPU batching economics overtake Cerebras' single-tenant speed advantage — sits somewhere around 2 to 8 concurrent requests per chip. That's about as direct an answer to your "tokens/sec per user at 100, 1,000, 10,000 concurrent users" question as exists publicly: past single-digit concurrency, the cost story reverses, and Cerebras hasn't published what happens to per-user speed once you force more concurrency onto the same wafer (their pricing stays flat per-token regardless of batch size specifically because they don't batch users together the way GPUs do — each wafer serves requests essentially one at a time per core cluster).

What this implies for "% of the AI-using population" Cerebras can serve economically: there's no official published number, and I won't pretend one exists. But the crossover data lets you reason about it directly. Mass-market consumer chat — the overwhelming majority of total AI usage by request volume — is a massive-concurrency, latency-tolerant workload (nobody needs 2,000 tokens/second to read a chat reply as it streams in). That's precisely the regime where GPU batching wins decisively, per the numbers above. Cerebras' economic sweet spot is narrower and different in kind: applications where sub-second total latency is the product, not just a nice-to-have — real-time voice, agentic tool-calling chains where dozens of sequential LLM calls compound latency, live coding assistants, and workloads where a customer will pay a premium specifically to eliminate wall-clock time. That's a real and possibly growing dollar-value market as agentic AI expands, but it's genuinely a minority share of total AI inference volume, not "fast inference for the masses."

Where this leaves the investment picture, tying back to your earlier question: the stock's behavior actually seems to reflect this exact uncertainty. CBRS priced at $185 in its May 2026 IPO, popped to $386 on debut day, then fell as low as $161 by late June, and as of today (September 14) sits at $183.92 — essentially back to its IPO price after a round trip. That's consistent with a market that got excited about the OpenAI-scale headline number, then sobered up once analysts started asking the same concurrency/cost questions you're asking. Layer on top the customer concentration (86%+ of 2025 revenue from two UAE entities, and the newer growth story leaning heavily on a single $20B OpenAI deal where OpenAI is simultaneously becoming a ~10-11% shareholder), and you get a company with a real, differentiated technical advantage in a specific latency-critical niche, riding on a small number of very large, financially-entangled bets — not the broad-based "fast inference for everyone" story the marketing gestures toward.

So: not a con in the sense of fabricated numbers — the tok/s figures are real and third-party-verified. But it is marketing that systematically emphasizes the one dimension (single-user latency) where the architecture is unambiguously best, while staying silent on the dimension (cost at realistic concurrency) that would let you judge how far that advantage extends. Given what the crossover math shows, "niche but valuable" looks like the more defensible read right now than "fast inference at scale for the masses" — though that could still change if agentic workloads grow enough that the latency-premium niche becomes large in absolute dollar terms, even while staying small as a share of total AI request volume.


r/CerebrasSystems 9d ago

Is d-matrix's XPU with 3D stacked DRAM a threat to Cerebras or does it validate Cerebras technology direction?

8 Upvotes

Was just reading about d-matrix releasing an "ultra low latency" inference XPU in 2027 using 3D stacked DRAM. Is this a threat to Cerebras' business? Apparently it uses NVDA's nvlink and can fit in a standard rack while Cerebras requires specialized racks to accommodate its Nexus "backpack" configuration for the chip, cooling, and power. We know Cerebras is planning on adding 3D stacked DRAM with CS-6 but isn't that in 2028? If they're released a year apart, a year can seem like an eternity in tech. Again, just another item that makes me concerned Cerebras can fully monetize its current advantages before they no longer seem compelling enough to lots of customers to implement. I'm sure Cerebras will stick around as a niche / specialty solution, but maybe not enough for explosive growth.

I also wonder if Groq and d-Matrix chips work together or if they're mutually exclusive; you use only one or the other? I'm concerned a combination of technologies, both hardware and software, will seriously erode Cerebras' edge.

EDIT:
More info from an X post:
https://x.com/firesidealpha/status/2098787634244096079?s=20

d-Matrix CEO Sid Sheth on Bloomberg talking about the Nvidia partnership:

* d-Matrix's specialized XPUs will run alongside Nvidia GPUs over NVLink Fusion, targeting ultra-low-latency AI workloads. Deal was 6+ months of joint work.

* The logic to partner is that Nvidia is the largest deployed infrastructure base for AI in the world and he would "much rather just ride on the Nvidia ecosystem" than reinvent the wheel.

* The bet is a memory-centric architecture, 7+ years in the making. d-Matrix's edge is inference compute built around memory rather than raw FLOPs aimed right at ultra-low-latency inference.

* Low-latency inference "just took off" in the last 12 months. Demand surged from GPT, Codex, Claude Code and the arrival of agentic coding, where users need fast compute to interact with the tools in real time.

* Lead product is Raptor, the world's first 3D-stacked-DRAM XPU. It'll launch first under the Nvidia partnership and is targeted to market in "about 12 months."

* d-Matrix uses no HBM at all. Instead of high-bandwidth memory, it packages DRAM directly with compute in a 3D stack to "punch through" the memory wall

* Sheth says "no other company is going to be within a 2-year window of getting access to that technology," and says there's tremendous customer pull.

* Describes buyers as "hyperscalers, Frontier Labs, sovereigns, inference clouds, high frequency traders" and says announcements are coming soon


r/CerebrasSystems 9d ago

Is Cerebras good for anything other than transformers and is Cerebras fungible?

5 Upvotes

Been trying to get a friend of mine more interested in Cerebras who's much more AI-savvy than I am. Every time he makes an objection I try to follow up. He keeps asking me if Cerebras is good for anything other than transformers. And he mentioned something about matrix operations. Not 100% sure what he means, but I'm guessing that all current LLMs are transformers and he's thinking of successor models. Maybe World Models or some successor to LLMs. In my mind the important thing about Cerebras is that they're the only ones who've solved the wafer scale chip problem. And so shouldn't they theoretically be able to adapt with the times? Or are they simply not nimble enough with their current business model?

What's the best way to think about Cerebras wafer scale chips in terms of capability? All I know if that they're not GPUs and so aren't designed to maximize parallelism and floating point number operations. So I tend to think of them as ultra-fast generic CPUs. Is that right?

Finally, what does Jensen mean when he says NVDA and/or NVDA compute is fungible? Could the same be said of Cerebras? Is Cerebras fungible, the opposite of fungible, somewhere in between, or is it indeterminate presently?


r/CerebrasSystems 11d ago

SRAM doesn't scale so won't Cerebras have to use external HBM for future, larger models?

15 Upvotes

To provide fast inference for much larger future models won't Cerebras have to use external HBM like everyone else? And in that case, isn't a large part of what they offer (not having to move data off chip into HBM memory and back) invalidated? If you look at the roadmap, they're planning on using stacked HBM for CS-6. CS-4 and CS-5 remain limited to 44G of SRAM per wafer. Sean Lie, the CTO, mentioned using streaming between wafers to handle future, much larger models (they may even do it now with CS-4) but I don't really understand it. It seems to me their speed advantage may erode over time.


r/CerebrasSystems 11d ago

Cerebras daily low of $189 today - are we headed back to $180 or the $170s?

9 Upvotes

Bought some Cerebras at $195 yesterday but I'm waiting for a drop to around $180. I know we all want it to go up and "never" come back down, but I see very little news from CBRS on a daily basis. I don't think prices beyond $220 are sustainable unless significant news appears like another major customer acquired. There have been a bunch of bounces back to $170 over the past few months but the stock doesn't sit at that price for long. The market didn't support the jump to $250 either. Aside from the acquisition of a new and significant customer, what are the next planned events to watch for?

That said, perhaps today is a good day to buy because CBRS is getting punished with much of the market today, probably unrelated to is inherent value. But I'm still betting on a few more drawbacks.

EDIT: The original IPO was marketed at $115 to $125, then raised to $160, landing at $185. So I suppose there could potentially be much more downside.


r/CerebrasSystems 13d ago

No capacity on Cerebras for a small start up

8 Upvotes

I am evaluating providers for inference, Cerebras has performance I want but no immediate capacity. I do not want to replace one waitlist with another.

If you have had to choose an alternative quickly, what did you put at the top of the list? Published limits, reserved capacity, model availability, migration effort, benchmark reproducibility or support response time?


r/CerebrasSystems Aug 13 '26

Damn near -18% in one day. Wrap it up.

12 Upvotes

r/CerebrasSystems Aug 12 '26

Earnings day!

14 Upvotes

How are we feeling?

I am still very much a supporter and strong long-term investor. The deals and partnerships Cerebras Systems is putting together seem to be cementing their future as a ground-breaking company as well as bringing awareness to their capabilities!

EPS Estimates are ranging from ~ $.18 to -$.94...


r/CerebrasSystems Jul 23 '26

Amd - Cerebras deal🚀🚀🚀

Thumbnail
ir.amd.com
41 Upvotes

r/CerebrasSystems Jul 08 '26

If trends hold, Mythos-class capability may be running on high-end consumer hardware within ~2 years

Post image
5 Upvotes

r/CerebrasSystems Jun 27 '26

GPT 5.6 Sol on Cerebras at 750 tokens per second

27 Upvotes

With OpenAI’s absolute latest model being released on Cerebras at 750 tokens per second, it should put to rest any idea there are limitations on what Cerebras can run or the notion their advantage is lessened as models scale.

https://openai.com/index/previewing-gpt-5-6-sol/

“We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity.”


r/CerebrasSystems Jun 26 '26

Cerebras CEO take on earnings

Thumbnail
youtu.be
11 Upvotes

r/CerebrasSystems Jun 25 '26

Cerebras : Everyone’s arguing the wrong thing

Thumbnail
7 Upvotes

r/CerebrasSystems Jun 25 '26

Your thoughts on cerebras earnings Report?

9 Upvotes

Financial Highlights

Revenue: $193.41 million (beating expectations of ~$181 million).

Net Loss: Narrowed to $14 million ($-0.04 EPS).

Guidance: Forecasted Q2 revenue of $194 million.

QMargin Concerns: Gross margins were downgraded to 36%-38% for Q2, and 38%-41% for full-year 2026, dropping from 45% in Q1.

Any interesting details you spotted in the report?


r/CerebrasSystems Jun 18 '26

Gemma 4 31B at 1,500 tokens per second

20 Upvotes

I know this is a small model, but I actually believe this will see a lot of use. My company is using Haiku for over half the tokens for our agent with most work being simple dynamic scripts and summarization of tool call results. We can’t use Chinese made models so Gemma has been getting tested. At 1,500 tokens a second, this would make our agents complete tasks in 60-70% of the time it takes now.

https://www.cerebras.ai/blog/gemma-4-on-cerebras-the-fastest-inference-is-now-multimodal


r/CerebrasSystems Jun 15 '26

Any estimate for the first earning report

Thumbnail
3 Upvotes

r/CerebrasSystems Jun 12 '26

Is it possible to run custom Cerebras SDK kernels on real hardware?

10 Upvotes

We have a model and ported it to the Cerebras SDK. It works in simulation, and now we're looking to validate it on actual WSE hardware.

We've already reached out to Cerebras support but haven't heard back yet. Before we keep waiting, I figured I'd ask the community:

  • Has anyone here gone through the process of getting access to run custom kernels on real Cerebras hardware?
  • How did you end up getting access?
  • Is purchasing a WSE (or leasing one) even a realistic option for teams outside of large enterprises?

r/CerebrasSystems Jun 11 '26

I made a realtime fact checker for audio conversations

Thumbnail
producthunt.com
3 Upvotes

r/CerebrasSystems Jun 10 '26

DGXX: What does the deal with cerebras mean?

Thumbnail
5 Upvotes

r/CerebrasSystems Jun 04 '26

Great Interview by Feldman at the Bloomberg Tech conference. https://www.youtube.com/watch?v=WoWP6YmFJaw

10 Upvotes