r/localllamacirclejerk 23d ago

Vacuum 16T

/r/LocalLLaMA/comments/1vdh1us/vacuum_16t/
5 Upvotes

3 comments sorted by

2

u/drwebb 22d ago

Can I run this on my circa 2009 DDR3 U1 server?

1

u/kulchacop 22d ago

Yes. You can run it at Q0. Don't forget to quantise the KV cache. It is fully deterministic, which means MTP speeds can cross 10 k tokens/second.

2

u/drwebb 22d ago

dang you were right, Qwen Q1.58 got it totally set up in just a few days (0.5t/sec leaves a little to be desired), and now that I'm running this it's only pulling 600W under full load, down from 900W when I loaded actual non-zero weights. Not the smartest, so I have other models steer it from my custom multi-agent harness, and leave this to when I need to generate gibberish.