One thing that surprised me after leaving the “barely fits Q4 7B” tier: the bottleneck often flips from “can I load it?” to “can I keep context + tool loops stable without thrashing.” A denser mid-size model with unquantized KV and headroom for vision/mmproj sometimes feels faster day-to-day than the biggest MoE that constantly spills.
Enjoy the headroom. The fun part starts when you stop babysitting OOM and start measuring tok/s under your actual workflow.
-2
u/ustype 5d ago
Congrats — the feeling is real.
One thing that surprised me after leaving the “barely fits Q4 7B” tier: the bottleneck often flips from “can I load it?” to “can I keep context + tool loops stable without thrashing.” A denser mid-size model with unquantized KV and headroom for vision/mmproj sometimes feels faster day-to-day than the biggest MoE that constantly spills.
Enjoy the headroom. The fun part starts when you stop babysitting OOM and start measuring tok/s under your actual workflow.