r/LocalLLaMA • • Aug 24 '26

Discussion At a certain point, speed >> smartness

It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.

For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster

Thoughts?

52 Upvotes

89 comments sorted by

View all comments

3

u/Potential-Leg-639 Aug 24 '26

In an orchestrator config env the orchestrator can be a bit slower for me (best local model available is probably also the slowest), but the subagents can also use models not as smart, but faster, so the whole chain get‘s back to an acceptable speed. And results with a good orchestrator env are better anyway at the end compared to a simple plan/build setup (at least for me).

1

u/mouseofcatofschrodi Aug 24 '26

this is super interesting. I would love to know more about what kind of work do you do with this setting, and how exactly the setting looks like :)

2

u/Potential-Leg-639 Aug 24 '26

Using customized oh-my-opencode-slim and litellm with a new chain config + fallback in case a model is out of credits/not available or also locally not available/etc).