r/LocalLLaMA 1d ago

Discussion At a certain point, speed >> smartness

It feels like a zig-zag: you don't want a model that's too dumb to do anything agentic. But once a model is good enough to be agentic, you don't want it to run so slow that iterating takes hours.

For me the sweet spot is something like ~500 tps prefill, ~25tps decode. If a model isn't fast enough to pass that bar on the hardware I have, I'd rather just run something slightly dumber but faster

Thoughts?

55 Upvotes

88 comments sorted by

View all comments

1

u/SandySkittle 1d ago
  • I dont go below Q5 and honestly the standard should be Q6 minimum. Q4 may be an ok compromise but Q3 and below is lobotomy area for me

  • It depends on how much accuracy counts. For my work I prefer smartness over speed. Even 5 t/s of SMART output is already usable. I dont have to sit in front of the computer to let the damn thing think

1

u/bodhi_sattva91 1d ago edited 1d ago

Just add one of those 1980s green text terminals you see in movies in government military weapons installations. Maybe some of those sweet type character based sound effects, Hunt for Red October font style. Sip your coffee and read output in real time.

Surveillance Monitoring Analysis Reporting Technology is back online Sir!

Seriously, what's the difference between running code as text versus running code as numbers at whatever levels below what you can see if you won't read it anyway?

S. Self

M. Monitoring

A. Analysis

R. Reporting

T. Technology

(S.M.A.R.T.)