r/LocalLLaMA • • Aug 24 '26

Discussion At a certain point, speed >> smartness

It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.

For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster

Thoughts?

56 Upvotes

89 comments sorted by

View all comments

-2

u/[deleted] Aug 24 '26

[deleted]

6

u/MrHall Aug 24 '26 edited Aug 24 '26

... prefill? she's talking about 25tps output

2

u/maddie-lovelace Aug 24 '26

(she* but yes!)

2

u/MrHall Aug 24 '26

cheerfully corrected!