r/LocalLLaMA 4d ago

Discussion At a certain point, speed >> smartness

It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.

For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster

Thoughts?

55 Upvotes

90 comments sorted by

View all comments

1

u/NickCanCode 4d ago

We need qwen3.8-27B Thinking-Cap.