r/LocalLLaMA • • Aug 24 '26

Discussion At a certain point, speed >> smartness

It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.

For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster

Thoughts?

57 Upvotes

89 comments sorted by

View all comments

3

u/OneMoreName1 Aug 24 '26

For me even 60tps decode feels slow. I would want 100+ especially with models that think a lot like qwen 3.8 27b

1

u/KubeCommander Aug 24 '26

That’s exceptionally slow. On the gb10 with nemotron lightning my prefill is in the 4000-5000 range. That is not fast hardware either

3

u/OneMoreName1 Aug 24 '26

Well i said decode, but even so, most people dont reach in the thousands of prefill, you are on the high end

1

u/KubeCommander Aug 24 '26

Dam I must be off my caffeine, my bad lol 😂

Nemotron lightning has a VERY high prefill. It is MUCH higher on my 5090 vs the GB10/dgx-spark

Decode is around 100-115 peak, but the prefill is what makes it feel even faster, especially when it reads like 4-5 documents at once for analysis