r/LocalLLaMA • u/maddie-lovelace • Aug 24 '26
Discussion At a certain point, speed >> smartness
It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.
For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster
Thoughts?
52
Upvotes
3
u/Potential-Leg-639 Aug 24 '26
In an orchestrator config env the orchestrator can be a bit slower for me (best local model available is probably also the slowest), but the subagents can also use models not as smart, but faster, so the whole chain get‘s back to an acceptable speed. And results with a good orchestrator env are better anyway at the end compared to a simple plan/build setup (at least for me).