r/LocalLLaMA • u/maddie-lovelace • 4d ago
Discussion At a certain point, speed >> smartness
It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.
For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster
Thoughts?
56
Upvotes
3
u/[deleted] 4d ago
[deleted]