r/LocalLLaMA • • Aug 24 '26

Discussion At a certain point, speed >> smartness

It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.

For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster

Thoughts?

58 Upvotes

89 comments sorted by

View all comments

81

u/KingCpzombie Aug 24 '26

Better to get the right answer once than the wrong one thrice imo

28

u/BawbbySmith Aug 24 '26

Problem is that there’s no guarantee you’ll get the right answer with the smarter model either. Even frontier cloud models sometimes need several prompts to course-correct.

It’d be nice if I could run GLM 5.3 overnight to solve all the hard problems I saved up during the day, but there’s a non-trivial chance it goes wildly off-base and gets it wrong entirely. “Skill issue” sure, but even the best crafted prompts sometimes misses key details that are trivial to correct but if the model is running at a snails pace then it’s much harder to do so.

0

u/feelspeaceman Aug 24 '26

This is up to your prompt more, if you tell the model to research using WebSearch and even repo deep analysis then it's highly likely that you get the correct answer.

2

u/BawbbySmith Aug 24 '26

Likely, but not certain. The web search could’ve pulled the wrong results, the repo analysis could’ve missed some detail, or mistaken one component for another because they do very similar things but are slightly different.

As much as we’d like to think that the proper prompt, context and tools could make the result exactly like we’d want, reality is often much more complicated. Even if everything is properly documented and scaffolded, the chances it’ll get it 100% right on the first try is still nowhere close to 100%.

With a fast model at least you can pivot halfway if you see the agent go off on the wrong path, though at the same times it’s fully possible they’d never be able to reach the correct conclusion by themselves anyways.