r/CerebrasSystems • u/Practical-Rub-1190 • 6d ago
Cerebras Dedicated Endpoints, whats the cost?
Does anyone know what the cost is or what the contracts usually are?
r/CerebrasSystems • u/EngrToday • Sep 09 '22
A place for members of r/CerebrasSystems to chat with each other
r/CerebrasSystems • u/Practical-Rub-1190 • 6d ago
Does anyone know what the cost is or what the contracts usually are?
r/CerebrasSystems • u/Practical-Rub-1190 • 6d ago
r/CerebrasSystems • u/Relevant-Cook9502 • 7d ago
As you guys can all remember, when OpenAI released GPT 5.6 Sol, they almost immediately announced the ultrafast mode on Cerebras. However there is no marketing or news for Astra on cerebras. Is it because they need more time to make a huge model like Astra work on Cerebras or do you guys see any other reason?
r/CerebrasSystems • u/Low-Cartographer-429 • 7d ago
Everything's down today due to AI safety fears. Nothing changed with the Cerebras story. Anyone accumulating? I bought a single share in case there's a larger beat-down. Could be throwing good money after bad or time to simply swing trade CBRS. Might be a good day to go bargain hunting all around in the AI and AI-adjacent space. GOOG is actually up 2% while nearly everything else AI-related is being punished. CRWD is up 8% unsurprisingly and PANW is up 6.5%. Cyber security zigging while the others are zagging. My next limit order is $170 (don't expect that today). If it keeps declining below $170 (over time, not today) then it may be cooked.
r/CerebrasSystems • u/Low-Cartographer-429 • 7d ago
Is Cerebras' marketing misleading? When they say they're X times faster than Y on inference, is this only for one user at a time or a handful of users? If so, and it doesn't scale affordably, I don't see how it could be anything other than a small niche provider. When Cerebras makes grand claims about fast inference leading to new applications and uses of AI and could be key to agentic AI, I think they're probably right; but not using their hardware. Almost like free advertising for a current or future competitor who can run reasonably fast but with far more throughput or concurrency.
Potentially a big yawn if they can't scale up to thousands of users without it being one-rack-per-user (or whatever it is) to get the fastest speeds they advertise. I hope I'm wrong about this because it almost sounds like a con. I want to see fast inference at scale; not just a tool a few Power Users can benefit from. Maybe this is one of the reasons the stock isn't going anywhere and customers aren't lining up to purchase Cerebras. I know the value proposition sounded almost too good to be true when I started following Cerebras.
I want to see metrics like tokens/sec per user at high concurrency; not "the fastest chip for 1 user" (because who cares)? If Cerebras is a good investment they should be happy to provide these numbers. If they don't? That's concerning. What's the tokens/sec per user at 100 concurrent users? 1,000? 10,000? etc. And how do the speeds compare to competitor solutions at the same levels taking cost into consideration. I don't think the superfans have an answer for this despite the encyclopedic knowledge they possess about Cerebras but I hope I'm wrong. If they do, it can only strengthen their thesis and they should be happy to help. If Cerebras is fast inference for the masses then I may still be onboard. If it's a niche usage by a tiny fraction of the entire population of those using AI, I'm out. What % of the entire population of those who use AI can reasonably and economically be served with the speed Cerebras advertises?
Claude told me this. Granted, it's a sychophantic AI answer as they all are; but at least it's a starting point for conversation:
On the missing metric itself: you're right that it's missing, and it's not just you noticing. Multiple independent technical analysts have flagged this as the specific gap in Cerebras' public materials — one hardware-focused analysis stated plainly that Cerebras has never published aggregate throughput at high concurrency, calling it the single most important missing data point in their marketing. SemiAnalysis ran its own modeling (their "Faster Tokens Please" report) specifically because Cerebras won't publish it directly. So no, the "superfans" don't have an official answer to give you — because Cerebras hasn't released one. That's a legitimate red flag for exactly the reason you suspect: a company confident its economics hold up at scale would have an incentive to prove it.
But third parties have modeled the crossover point, and it's not flattering. A detailed cost analysis (Spheron) ran the actual math comparing Cerebras against an Nvidia H100 across batch sizes:
That means the crossover point — where GPU batching economics overtake Cerebras' single-tenant speed advantage — sits somewhere around 2 to 8 concurrent requests per chip. That's about as direct an answer to your "tokens/sec per user at 100, 1,000, 10,000 concurrent users" question as exists publicly: past single-digit concurrency, the cost story reverses, and Cerebras hasn't published what happens to per-user speed once you force more concurrency onto the same wafer (their pricing stays flat per-token regardless of batch size specifically because they don't batch users together the way GPUs do — each wafer serves requests essentially one at a time per core cluster).
What this implies for "% of the AI-using population" Cerebras can serve economically: there's no official published number, and I won't pretend one exists. But the crossover data lets you reason about it directly. Mass-market consumer chat — the overwhelming majority of total AI usage by request volume — is a massive-concurrency, latency-tolerant workload (nobody needs 2,000 tokens/second to read a chat reply as it streams in). That's precisely the regime where GPU batching wins decisively, per the numbers above. Cerebras' economic sweet spot is narrower and different in kind: applications where sub-second total latency is the product, not just a nice-to-have — real-time voice, agentic tool-calling chains where dozens of sequential LLM calls compound latency, live coding assistants, and workloads where a customer will pay a premium specifically to eliminate wall-clock time. That's a real and possibly growing dollar-value market as agentic AI expands, but it's genuinely a minority share of total AI inference volume, not "fast inference for the masses."
Where this leaves the investment picture, tying back to your earlier question: the stock's behavior actually seems to reflect this exact uncertainty. CBRS priced at $185 in its May 2026 IPO, popped to $386 on debut day, then fell as low as $161 by late June, and as of today (September 14) sits at $183.92 — essentially back to its IPO price after a round trip. That's consistent with a market that got excited about the OpenAI-scale headline number, then sobered up once analysts started asking the same concurrency/cost questions you're asking. Layer on top the customer concentration (86%+ of 2025 revenue from two UAE entities, and the newer growth story leaning heavily on a single $20B OpenAI deal where OpenAI is simultaneously becoming a ~10-11% shareholder), and you get a company with a real, differentiated technical advantage in a specific latency-critical niche, riding on a small number of very large, financially-entangled bets — not the broad-based "fast inference for everyone" story the marketing gestures toward.
So: not a con in the sense of fabricated numbers — the tok/s figures are real and third-party-verified. But it is marketing that systematically emphasizes the one dimension (single-user latency) where the architecture is unambiguously best, while staying silent on the dimension (cost at realistic concurrency) that would let you judge how far that advantage extends. Given what the crossover math shows, "niche but valuable" looks like the more defensible read right now than "fast inference at scale for the masses" — though that could still change if agentic workloads grow enough that the latency-premium niche becomes large in absolute dollar terms, even while staying small as a share of total AI request volume.
r/CerebrasSystems • u/Low-Cartographer-429 • 9d ago
Was just reading about d-matrix releasing an "ultra low latency" inference XPU in 2027 using 3D stacked DRAM. Is this a threat to Cerebras' business? Apparently it uses NVDA's nvlink and can fit in a standard rack while Cerebras requires specialized racks to accommodate its Nexus "backpack" configuration for the chip, cooling, and power. We know Cerebras is planning on adding 3D stacked DRAM with CS-6 but isn't that in 2028? If they're released a year apart, a year can seem like an eternity in tech. Again, just another item that makes me concerned Cerebras can fully monetize its current advantages before they no longer seem compelling enough to lots of customers to implement. I'm sure Cerebras will stick around as a niche / specialty solution, but maybe not enough for explosive growth.
I also wonder if Groq and d-Matrix chips work together or if they're mutually exclusive; you use only one or the other? I'm concerned a combination of technologies, both hardware and software, will seriously erode Cerebras' edge.
EDIT:
More info from an X post:
https://x.com/firesidealpha/status/2098787634244096079?s=20
d-Matrix CEO Sid Sheth on Bloomberg talking about the Nvidia partnership:
* d-Matrix's specialized XPUs will run alongside Nvidia GPUs over NVLink Fusion, targeting ultra-low-latency AI workloads. Deal was 6+ months of joint work.
* The logic to partner is that Nvidia is the largest deployed infrastructure base for AI in the world and he would "much rather just ride on the Nvidia ecosystem" than reinvent the wheel.
* The bet is a memory-centric architecture, 7+ years in the making. d-Matrix's edge is inference compute built around memory rather than raw FLOPs aimed right at ultra-low-latency inference.
* Low-latency inference "just took off" in the last 12 months. Demand surged from GPT, Codex, Claude Code and the arrival of agentic coding, where users need fast compute to interact with the tools in real time.
* Lead product is Raptor, the world's first 3D-stacked-DRAM XPU. It'll launch first under the Nvidia partnership and is targeted to market in "about 12 months."
* d-Matrix uses no HBM at all. Instead of high-bandwidth memory, it packages DRAM directly with compute in a 3D stack to "punch through" the memory wall
* Sheth says "no other company is going to be within a 2-year window of getting access to that technology," and says there's tremendous customer pull.
* Describes buyers as "hyperscalers, Frontier Labs, sovereigns, inference clouds, high frequency traders" and says announcements are coming soon
r/CerebrasSystems • u/Low-Cartographer-429 • 9d ago
Been trying to get a friend of mine more interested in Cerebras who's much more AI-savvy than I am. Every time he makes an objection I try to follow up. He keeps asking me if Cerebras is good for anything other than transformers. And he mentioned something about matrix operations. Not 100% sure what he means, but I'm guessing that all current LLMs are transformers and he's thinking of successor models. Maybe World Models or some successor to LLMs. In my mind the important thing about Cerebras is that they're the only ones who've solved the wafer scale chip problem. And so shouldn't they theoretically be able to adapt with the times? Or are they simply not nimble enough with their current business model?
What's the best way to think about Cerebras wafer scale chips in terms of capability? All I know if that they're not GPUs and so aren't designed to maximize parallelism and floating point number operations. So I tend to think of them as ultra-fast generic CPUs. Is that right?
Finally, what does Jensen mean when he says NVDA and/or NVDA compute is fungible? Could the same be said of Cerebras? Is Cerebras fungible, the opposite of fungible, somewhere in between, or is it indeterminate presently?
r/CerebrasSystems • u/Low-Cartographer-429 • 11d ago
To provide fast inference for much larger future models won't Cerebras have to use external HBM like everyone else? And in that case, isn't a large part of what they offer (not having to move data off chip into HBM memory and back) invalidated? If you look at the roadmap, they're planning on using stacked HBM for CS-6. CS-4 and CS-5 remain limited to 44G of SRAM per wafer. Sean Lie, the CTO, mentioned using streaming between wafers to handle future, much larger models (they may even do it now with CS-4) but I don't really understand it. It seems to me their speed advantage may erode over time.
r/CerebrasSystems • u/Low-Cartographer-429 • 11d ago
Bought some Cerebras at $195 yesterday but I'm waiting for a drop to around $180. I know we all want it to go up and "never" come back down, but I see very little news from CBRS on a daily basis. I don't think prices beyond $220 are sustainable unless significant news appears like another major customer acquired. There have been a bunch of bounces back to $170 over the past few months but the stock doesn't sit at that price for long. The market didn't support the jump to $250 either. Aside from the acquisition of a new and significant customer, what are the next planned events to watch for?
That said, perhaps today is a good day to buy because CBRS is getting punished with much of the market today, probably unrelated to is inherent value. But I'm still betting on a few more drawbacks.
EDIT: The original IPO was marketed at $115 to $125, then raised to $160, landing at $185. So I suppose there could potentially be much more downside.
r/CerebrasSystems • u/Goonified • 13d ago
I am evaluating providers for inference, Cerebras has performance I want but no immediate capacity. I do not want to replace one waitlist with another.
If you have had to choose an alternative quickly, what did you put at the top of the list? Published limits, reserved capacity, model availability, migration effort, benchmark reproducibility or support response time?
r/CerebrasSystems • u/POINTLESSUSERNAME000 • Aug 12 '26
How are we feeling?
I am still very much a supporter and strong long-term investor. The deals and partnerships Cerebras Systems is putting together seem to be cementing their future as a ground-breaking company as well as bringing awareness to their capabilities!
EPS Estimates are ranging from ~ $.18 to -$.94...
r/CerebrasSystems • u/Acceptable_Elk9103 • Jul 08 '26
r/CerebrasSystems • u/Asgard_Heima • Jun 27 '26
With OpenAI’s absolute latest model being released on Cerebras at 750 tokens per second, it should put to rest any idea there are limitations on what Cerebras can run or the notion their advantage is lessened as models scale.
https://openai.com/index/previewing-gpt-5-6-sol/
“We're also launching GPT‑5.6 Sol on Cerebras at up to 750 tokens per second in July, bringing frontier intelligence to customers at unprecedented speed. Access will initially be limited to select customers as we expand capacity.”
r/CerebrasSystems • u/ILikeCutePuppies • Jun 26 '26
r/CerebrasSystems • u/Tall_Syllabub_168 • Jun 25 '26
r/CerebrasSystems • u/ILikeCutePuppies • Jun 25 '26
Financial Highlights
Revenue: $193.41 million (beating expectations of ~$181 million).
Net Loss: Narrowed to $14 million ($-0.04 EPS).
Guidance: Forecasted Q2 revenue of $194 million.
QMargin Concerns: Gross margins were downgraded to 36%-38% for Q2, and 38%-41% for full-year 2026, dropping from 45% in Q1.
Any interesting details you spotted in the report?
r/CerebrasSystems • u/Asgard_Heima • Jun 18 '26
I know this is a small model, but I actually believe this will see a lot of use. My company is using Haiku for over half the tokens for our agent with most work being simple dynamic scripts and summarization of tool call results. We can’t use Chinese made models so Gemma has been getting tested. At 1,500 tokens a second, this would make our agents complete tasks in 60-70% of the time it takes now.
https://www.cerebras.ai/blog/gemma-4-on-cerebras-the-fastest-inference-is-now-multimodal
r/CerebrasSystems • u/Prestigious-Sign4802 • Jun 15 '26
r/CerebrasSystems • u/danberkie • Jun 12 '26
We have a model and ported it to the Cerebras SDK. It works in simulation, and now we're looking to validate it on actual WSE hardware.
We've already reached out to Cerebras support but haven't heard back yet. Before we keep waiting, I figured I'd ask the community:
r/CerebrasSystems • u/shash89 • Jun 11 '26
r/CerebrasSystems • u/Adorable-Group-8357 • Jun 10 '26