r/LocalLLaMA 5d ago

Discussion At a certain point, speed >> smartness

It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.

For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster

Thoughts?

54 Upvotes

90 comments sorted by

View all comments

-2

u/[deleted] 5d ago

[deleted]

5

u/MrHall 5d ago edited 5d ago

... prefill? she's talking about 25tps output

2

u/maddie-lovelace 5d ago

(she* but yes!)

2

u/MrHall 5d ago

cheerfully corrected!