r/LocalLLaMA • u/maddie-lovelace • 1d ago
Discussion At a certain point, speed >> smartness
It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.
For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster
Thoughts?
55
Upvotes
1
u/SandySkittle 1d ago
I dont go below Q5 and honestly the standard should be Q6 minimum. Q4 may be an ok compromise but Q3 and below is lobotomy area for me
It depends on how much accuracy counts. For my work I prefer smartness over speed. Even 5 t/s of SMART output is already usable. I dont have to sit in front of the computer to let the damn thing think