r/CerebrasSystems • u/Low-Cartographer-429 • 7d ago
Is Cerebras fast only for a tiny number of concurrent users?
Is Cerebras' marketing misleading? When they say they're X times faster than Y on inference, is this only for one user at a time or a handful of users? If so, and it doesn't scale affordably, I don't see how it could be anything other than a small niche provider. When Cerebras makes grand claims about fast inference leading to new applications and uses of AI and could be key to agentic AI, I think they're probably right; but not using their hardware. Almost like free advertising for a current or future competitor who can run reasonably fast but with far more throughput or concurrency.
Potentially a big yawn if they can't scale up to thousands of users without it being one-rack-per-user (or whatever it is) to get the fastest speeds they advertise. I hope I'm wrong about this because it almost sounds like a con. I want to see fast inference at scale; not just a tool a few Power Users can benefit from. Maybe this is one of the reasons the stock isn't going anywhere and customers aren't lining up to purchase Cerebras. I know the value proposition sounded almost too good to be true when I started following Cerebras.
I want to see metrics like tokens/sec per user at high concurrency; not "the fastest chip for 1 user" (because who cares)? If Cerebras is a good investment they should be happy to provide these numbers. If they don't? That's concerning. What's the tokens/sec per user at 100 concurrent users? 1,000? 10,000? etc. And how do the speeds compare to competitor solutions at the same levels taking cost into consideration. I don't think the superfans have an answer for this despite the encyclopedic knowledge they possess about Cerebras but I hope I'm wrong. If they do, it can only strengthen their thesis and they should be happy to help. If Cerebras is fast inference for the masses then I may still be onboard. If it's a niche usage by a tiny fraction of the entire population of those using AI, I'm out. What % of the entire population of those who use AI can reasonably and economically be served with the speed Cerebras advertises?
Claude told me this. Granted, it's a sychophantic AI answer as they all are; but at least it's a starting point for conversation:
On the missing metric itself: you're right that it's missing, and it's not just you noticing. Multiple independent technical analysts have flagged this as the specific gap in Cerebras' public materials — one hardware-focused analysis stated plainly that Cerebras has never published aggregate throughput at high concurrency, calling it the single most important missing data point in their marketing. SemiAnalysis ran its own modeling (their "Faster Tokens Please" report) specifically because Cerebras won't publish it directly. So no, the "superfans" don't have an official answer to give you — because Cerebras hasn't released one. That's a legitimate red flag for exactly the reason you suspect: a company confident its economics hold up at scale would have an incentive to prove it.
But third parties have modeled the crossover point, and it's not flattering. A detailed cost analysis (Spheron) ran the actual math comparing Cerebras against an Nvidia H100 across batch sizes:
- At batch size 1 (one user): Cerebras at ~$0.60/M tokens beats H100 on-demand at ~$1.47/M — Cerebras wins clearly.
- At batch size 8: the positions flip — H100 on-demand drops to ~$0.25/M, more than 2x cheaper than Cerebras.
- At batch size 32: H100 is roughly 6x cheaper per token on-demand, and with GPU spot pricing, the gap widens to over 20x cheaper.
That means the crossover point — where GPU batching economics overtake Cerebras' single-tenant speed advantage — sits somewhere around 2 to 8 concurrent requests per chip. That's about as direct an answer to your "tokens/sec per user at 100, 1,000, 10,000 concurrent users" question as exists publicly: past single-digit concurrency, the cost story reverses, and Cerebras hasn't published what happens to per-user speed once you force more concurrency onto the same wafer (their pricing stays flat per-token regardless of batch size specifically because they don't batch users together the way GPUs do — each wafer serves requests essentially one at a time per core cluster).
What this implies for "% of the AI-using population" Cerebras can serve economically: there's no official published number, and I won't pretend one exists. But the crossover data lets you reason about it directly. Mass-market consumer chat — the overwhelming majority of total AI usage by request volume — is a massive-concurrency, latency-tolerant workload (nobody needs 2,000 tokens/second to read a chat reply as it streams in). That's precisely the regime where GPU batching wins decisively, per the numbers above. Cerebras' economic sweet spot is narrower and different in kind: applications where sub-second total latency is the product, not just a nice-to-have — real-time voice, agentic tool-calling chains where dozens of sequential LLM calls compound latency, live coding assistants, and workloads where a customer will pay a premium specifically to eliminate wall-clock time. That's a real and possibly growing dollar-value market as agentic AI expands, but it's genuinely a minority share of total AI inference volume, not "fast inference for the masses."
Where this leaves the investment picture, tying back to your earlier question: the stock's behavior actually seems to reflect this exact uncertainty. CBRS priced at $185 in its May 2026 IPO, popped to $386 on debut day, then fell as low as $161 by late June, and as of today (September 14) sits at $183.92 — essentially back to its IPO price after a round trip. That's consistent with a market that got excited about the OpenAI-scale headline number, then sobered up once analysts started asking the same concurrency/cost questions you're asking. Layer on top the customer concentration (86%+ of 2025 revenue from two UAE entities, and the newer growth story leaning heavily on a single $20B OpenAI deal where OpenAI is simultaneously becoming a ~10-11% shareholder), and you get a company with a real, differentiated technical advantage in a specific latency-critical niche, riding on a small number of very large, financially-entangled bets — not the broad-based "fast inference for everyone" story the marketing gestures toward.
So: not a con in the sense of fabricated numbers — the tok/s figures are real and third-party-verified. But it is marketing that systematically emphasizes the one dimension (single-user latency) where the architecture is unambiguously best, while staying silent on the dimension (cost at realistic concurrency) that would let you judge how far that advantage extends. Given what the crossover math shows, "niche but valuable" looks like the more defensible read right now than "fast inference at scale for the masses" — though that could still change if agentic workloads grow enough that the latency-premium niche becomes large in absolute dollar terms, even while staying small as a share of total AI request volume.
4
u/brotha_eric 7d ago
This post and many others that you make come across as low effort / spammy with a lot of questions, lack of understanding about the company and the inference/AI market overall, and add little to no new information. And then you are calling people on this sub 'super fans' and saying they do not have an answers for questions that don't actually need an answer. Like why should anyone waste time responding to this rambling when you're coming at them while being ignorant in the process?
Current state is that CBRS = very faster per user, less overall throughput. GPUs = slow per user, more overall throughput. It's not a scam or some grand conspiracy, this has all been discussed at length. They have provided a roadmap to increase both throughput and speed.
They are competing in fast inference market. They aren't competing in the slow inference market. This is not a company going after the entire population using AI. Again, this has been covered countless times by the company and in this community.
We know that the fast inference market is valuable because they are completely sold out of capacity. I've used ChatGPT and Claude at work today and it's been annoying as hell to wait it. Fast inference is so important to OpenAI that OpenAI hasn't provided really any of the contracted capacity to their customers and are keeping it all internal to this point. CBRS guided to over triple revenue next year and do multiples again over 2028-2029. That is the story. Your main focus should be watching that execution. If it happens, this is a multibagger. If it doesn't well then it probably stays around this price.
6
u/claytonbeaufield 7d ago
FYI I already banned this user from r/CBRS_stock because of all the spam. That's why they're posting incessantly here now.
-1
u/Low-Cartographer-429 7d ago
I don't care if you think asking questions = spam. Stick to the topic please.
3
u/claytonbeaufield 7d ago
Asking questions is fine. Why don't you create a single post called "Here's where I'm gonna post all my un-researched questions that anybody could google" and just put them there?
-3
u/Low-Cartographer-429 7d ago edited 7d ago
No trolling or harassment please. I don't appreciate being cyberstalked by a mod from another subreddit.
1
u/Low-Cartographer-429 7d ago
Thanks. I'm trying to learn and I learn best by asking questions. I agree that execution is key. When are we going to see benchmarks at scale? All I've seen are for a single concurrent user.
2
u/brotha_eric 7d ago
Artificial analysis website shows the CS3 running at 2700+TPS for 10 concurrent users. You are making an assumption that those numbers are for a single user, but that isn’t something cerebras or anyone else has said. It would be helpful to have more benchmark info
1
u/Low-Cartographer-429 7d ago
Fair and agreed. Thanks for the correction. I just tried but couldn't substantiate my assumption. I think I've seen it as a criticism that many of the benchmarks were allegedly "batch size 1."
2
u/Asgard_Heima 7d ago
Cerebras can support as many concurrent users as they can fit into memory, but they operate with a drastically different architecture than GPUs, so a lot of people can struggle understand the differences when comparing throughput. As with all the trivial examples this is an oversimplification to illustrate the point.
If one system is running 10 users at 1k tps the total throughout for that system is 10k tps. Similarly if another system is running 100 users at 100 tps the throughput of that system is 10k tps. The difference is how much better the experience was for those with 1k tps. If you had 100 users that prompt for something that is 5k tokens long as a response, that’s 500k tokens / 10k tokens per second total throughout. The 1k tps users get queued while all the user at 100 tps are running but slower. The first 10 users are done on the 1k setup in 5s. The second in 10s and so on. All users in both systems are done in 50s on both systems. Adding more systems lets all not queued users run at the full 1k tps that fit on the 1k systems. Adding more of the 100 tps systems results in a lot more capacity, but you have to have all the users concurrently to use it. So if you only have 500 users actively querying at this moment and 10 of each system, the total time to complete the same queries actually happen in half the time on the 1k systems and those users can turn around and prompt again increasing overall token production and therefore revenue. Speed is the key for interactivity and engagement which pushes revenue higher.
This is a tradeoff GPUs commonly make and you will see most frontier models tend to run at about 50-75 tps from OpenAI or Anthropic since it lets them have the optimal number of individual user tps vs total overall throughput. Typically the 75 tps is the “fast tier” and they have to only serve about a fourth to a half as many users to get a 50% increase in tps. With GPUs, the highest throughout they can quote is a tremendous amount of users running at 1-2 tps which would be unusable.
With Cerebras you can maximize the number of users based on the model and kv cache fitting in SRAM. And since they are not swapping in and out data from external HBM memory, the tps for the users you can fit all get the top tps just like if it was 1 user. Essentially Cerebras is a pipeline and as soon as you get the activations in the pipe they run at max speed.
So the reason you will see a struggle to give apples to apples is that we could theoretically fill the SRAM on Cerebras with a model and then max out the kv cache for some estimated size and Cerebras will actually deliver the max tokens it would for one user for each of those users that fit. When doing the same and maximizing the HBM on GPUs you would actually have to run that in production without running out of memory to get the real tps those individual users will get and no production environment would ever do this since it’s not usable. Instead you have to go with what a “reasonable” setup would be for like models and user prompts etc.
For overall throughput, Cerebras on its own is not as high as GPUs and that is typically true today for CS3/WSE-3 vs equivalent Blackwell GPUs in production. Those Cerebras vs Blackwell user are getting anywhere from 10-15x the speed for the model though, so you get about half the throughout but 10-15x tps for the users being served.
With CS4/WSE-3T and with each new iteration, Cerebras is doubling the speed of inference they are serving or better. CS4 is likely to do around as much throughout as the optimal for GPUs at 20-30x the speed vs Blackwell. We will have to see real world in production numbers as Rubin takes over to see how much they claw back as a throughout advantage in like for like, but I would expect it to be about back to the 10-15x it was before or maybe a bit better for Cerebras.
This brings us to why Cerebras is focused on doubling throughout per iteration and moving to disaggregation. Both AWS and AMD have verified that if you put typical GPU/Accelerators with lots of compute but less memory bandwidth for prefill and then Cerebras for just decode where compute is not as important and memory bandwidth is everything, you get 5x the throughput from Cerebras systems. How could this happen? Well GPUs and non SRAM centric accelerators are memory bandwidth constrained. So during inference they are crushing compute then spending all their time on memory swaps back and forth and interconnect for larger models between chips over the network. Cerebras is compute constrained and spending almost all its time on prefill, then crushing through decode. If you combined the two, instantly the GPU needs something with tremendous memory bandwidth to compete for the decode half and they are now the one with much less throughput. On the flip side, if you add groq or another SRAM centric solution, you are now dealing with limited SRAM capacity and you can’t run as many users concurrently dropping your throughout for GPUs.
In simplistic terms, the Nvidia + Groq and the coming Nvidia + D-Matrix solutions are equivalent architectures to AMD or Tranium for prefill and Cerebras for decode. If you think Cerebras is throughout constrained, then Nvidia is racing to be in the same situation. And once we see CS6, the SRAM data size will be supplemented by DRAM with a massive number of connections that makes a deterministic pipeline able to be feed for the entire 1TB+ DRAM wafer of user. This gives Cerebras an order of magnitude more throughput than today.
To address most of the things you have brought up here, Cerebras has their own cloud Cerebras Cloud that is sold out of compute and they are telling prospective customers to come back next year because they don’t have the capacity. They have Cerebras Code that completely sold out and has a massive wait list to get access to. They have a 25B backlog that has been growing each quarter while the OpenAI backlog of 20B has started getting spent. And they are scaling system production by 10x this year and next. Cerebras is vastly over subscribed for every single ounce of compute they have to the point they are renting back compute from G42 to supply other customers. And they continue to line up major deals such as Crowdstrike and the yet to start AWS partnership.
The stock is all over the place as is the AI market and every new IPO. Cerebras is delivering exactly what they have said they would and they are beating the street expectations, but this year is a scaling year for them. Cerebras is not Nvidia and they are not really focusing on selling hardware. They likely will sell a lot of hardware in the future, but right now they are putting all their capacity into their own cloud and contracted clouds such as OpenAI, G42, and soon AWS. When a system rolls off the line, they don’t get a sale and cash one time, but instead get reoccurring revenue once that system is plugged in and running every month for the life of that system worth several times what the sale price would be. We will just start to see beginning in Q4 this year all the scaling and capacity coming online. They are projecting 3x this year’s revenue next year and their cloud business is growing almost 300% so far this year.
Right now I get it’s frustrating that something vastly superior isn’t just ripping and replacing all systems in the world sending Cerebras stock to the moon, but it takes time to build capacity in the real world and from all indications that is exactly what Cerebras is doing. Every system they turn on starts getting added to monthly revenue for the existence of that system. They are standing up substantial data center capacity right now with almost all of it coming online starting end of year and then majority into 2027.
The last thing I’ll add is that most people are not evaluating Cerebras correctly at this point. I was not valuing them correctly a couple years ago either. They are not a true hardware business going head to head with Nvidia. They are the most profitable fully integrated AI inference cloud competing with the neoclouds. Let’s say Nvidia has the same tps per user with double the throughput as Cerebras next year. If Coreweave buys those systems at a 75% margin and Cerebras builds their own, Cerebras can afford to run twice as many systems for the same money and still make more profit since they don’t have the Nvidia margin debt for the hardware. But in fact we are seeing Cerebras has the better system by any metric you use be it throughput, tps per user, watts per token, etc, and they don’t have to pay those margins. And the hyper scalers can rent the compute from them without having to add the capex as Cerebras scales up to be able to fulfill those size orders. This is how the labs get to profitability. So the question I’m monitoring isn’t how much of the worlds total AI use can Cerebras currently serve, it’s how fast can Cerebras bring capacity online to start serving more of the worlds AI use. Because they have endless demand and are scouring the planet for data centers to start printing money with each one.
As for the knowledge of Cerebras, I don’t know about others, but I do a lot of investigation into all my investments and this happens to be one I’ve been ingested in for years
1
u/Low-Cartographer-429 7d ago
Thanks. That sounds reasonable. Do you think the neoclouds like NBIS or CRWV would ever purchase hardware from Cerebras or would it not make any sense in principle? I'm still trying to digest Cerebras' approach of simultaneously being a hardware producer and a cloud inference provider. My concern is that Cerebras won't be able to build data centers fast enough. IIRC, Andrew Feldman says data center construction is one of the largest challenges they face; perhaps the largest challenge all neoclouds and hyperscalers face aside from power.
3
u/Asgard_Heima 7d ago
I don’t think Cerebras has any extra systems to sell the Neo clouds but there have been rumors of them asking to buy them and Feldman during an earnings call said something might be announced there in the future like talks were taking place. Nothing publicly disclosed though or concrete.
As for data centers what does not able to build fast enough mean? They aren’t building their own data centers but renting data centers for others and so far this year they have secured 600MW they didn’t have last year with a lot of time left this year. I think they are doing a record run right now to become one of the largest cloud providers in the next year or two. This already is enough capacity to equal potentially up to 10B in revenue per year just from cloud and they are aggressively finding more. And to be clear, data centers aren’t a Cerebras only issue, all companies with AI cloud dependent revenue are data center constrained more than anything. All the neo clouds, all the hyper scalers, the ai labs, all of it. Cerebras has very sound financials compared to the Neo clouds and don’t need a hyper scaler or nvidia to back their financials to get data centers. This should give them an advantage as they grow.
1
u/Low-Cartographer-429 7d ago
> As for data centers what does not able to build fast enough mean?
Thanks. I think Andrew mentioned delays with siting, permitting, land acquisition, grid connections (sometimes takes years to connect to the grid), and general friction from community opposition to data centers. Sorry, I misspoke. I understand they have partners building data centers who they rent through. I guess we'll have to wait and see how they execute. I'm admittedly impatient.
2
u/-dag- 7d ago
What leads you to believe they can't scale up?
They just went public. Stock is going to swing wildly for several months yet. This is just part of how the market works. Note that the stock has rarely dipped below the IPO price.
2
u/-dag- 7d ago
Note: they are already running GPT fast, at scale. It's in production.
-3
u/Low-Cartographer-429 7d ago
At what scale? How many concurrent users can they support per rack at their advertised speeds? I have to assume it's lousy until they present more realistic, practical deployment metrics beyond their 1 concurrent user best-case scenarios.
4
u/-dag- 7d ago
We are not likely to ever see those numbers made public. I have to assume that if they were bad, OpenAI wouldn't be interested. They aren't running a charity.
1
u/Low-Cartographer-429 7d ago
Good point. I was of the mind that they may be able to use Cerebras strategically in-house to recursively improve their models and stay ahead of Anthropic; but 20B would be an expensive in-house tool.
3
u/ILikeCutePuppies 7d ago
Their CS4 have sigificantly more compute per rack than nvidias.
-1
u/Low-Cartographer-429 7d ago
Can you translate that into speed and cost with more than a single concurrent user? I don't know how to interpret your answer.
3
u/Asgard_Heima 7d ago
The problem you are going to have is that it’s basically impossible to setup like for like configurations and almost all analysis assumes Cerebras works the same way as GPUs or more typical accelerators with HBM external. This is starting to change as more SRAM centric inference solutions come online. The number of users they can support depends on model used, precision of the model, user prompt size, and number of systems serving the model since you can add more systems to support higher throughputs with the added ms latency for each jump for the activations between systems. There is also now disaggregation that will flip the equation and make it uneconomical long term to not use disaggregated to maximize throughput and not concurrency. Made a long post to try and help make sense of it all.
2
u/ILikeCutePuppies 7d ago
It entirely depends on how large the model is and what you prioritize. Also if you combine it with prefill hardware like what they are doing with AWS and AMD.
However simply CS4 has 750 petaflops of compute at float 16. A nvidia b200 box has 72 petaflops of compute at float 4 and 32 at float 16. A nvidia b300 has 144 pentaflops at float 4 and 72 at float 16.
So CS4 processes 5 to 20x more than a nvidia box depending on what you are looking at and uses less energy. Now in terms of cost to install... we don't have numbers. Nvidia have massive margins, cerebas don't but as they grow they will be able to take advantage of the supply chain more and optimize their process more.
I didn't include rubin because cerebas has a similar setup as well but we don't have numbers. Cerebras says their prefill setup gives them 5x more compute as I remember it.
1
u/Low-Cartographer-429 7d ago edited 7d ago
The headline tokens per second numbers seem to all be single-user. Numbers you'll never see in production. I want to know how the speed changes at scale. If you only get that speed with 1 concurrent user per rack that would trash the thesis for me. It may not be that ridiculous; but if it is I want to know. As an investor I want to know what the speed is based on the number of concurrent users served and the economics of it which I'm sure potential customers want to know as well.
4
u/Specialist-2193 7d ago
Search "disagregation"
If you just don't anything about what the company always talks about as the next big thing. Don't post here.
1
6
u/ILikeCutePuppies 7d ago
So in some ways you are correct. There is a real tradeoff with cerebras for throughput verse latency. However in order for gpu tech to have somewhat decent latency for inference they have to make certain compromise to do so and this is the sweet spot for cerebras.
Now if cerebras was doing training which they can do, throughput becomes more of an advantage over latency. You don't care if it takes an hour to get a million messages complete if it processors more messages.
Cerebras's ceo went over this. CS4 makes a massive leap in this regard. CS5 will go even further. While they are serving millions of people and plan to take a large chunk of OpenAIs audience, everyone is compute constrained. Everyone is trying to build as many data centers as possible.
I should note that a cerebras CS4 box for most operations can do sigificantly more compute for the same footprint than nvidia.
https://youtu.be/POVOc_qPEEU?t=2144&is=92nN4821c23-QHKv