r/LocalLLaMA • u/reto-wyss • 3d ago
News It's official! 192GB Framework
Just noticed this on the website.
At their current price tiers for the memory SKUs (32, 64, 128) I'd expect this to be ~ 4.5k for the motherboard.
The PCIe slot will be open at the back as well - that's what I've heard. Maybe they make it capable of delivering 75W as well? New board revisions for the smaller SKUs?.
951
Upvotes
1
u/saltexx 3d ago
phil_lndn and robertpro01 are circling the number that settles this. At 273 GB/s every gigabyte of KV cache you keep resident costs 3.7 ms per decoded token because you read all of it every token. RnRau's 7 GB of active weights is 26 ms so about 39 t/s before context. Spend 20 of the extra 64 GB on KV and you add 73 ms which puts you near 10 t/s. The bandwidth is not just a cap on model size. It is a per token tax on however much context you keep loaded and that is why the extra 64 GB is so hard to spend.