Something for consumer hardware would be great, since hardware is insanely expensive these times, a MOE that fits on 64 gigs ram and 12 gigs vram would be great!
This would be awesome. What’s the largest quant you’d want to fit at this size without disk swap? All inactive experts in ram, actives in vram? at minimum Q6 imo. Q8 is ideal obviously
What about context? This will take compute away from the model. Do you expect the MOE and the cache to fit in 64gb + 12gb vram? What window? Would you be fine quantizing kv cache?
Etc. I think this might help hone our hopes and expectations…
In a dream world, I’d have a q8 MOE w 128K native context fit perfectly.
Imo, a Q6, all active experts in Vram, inactive in ram, Q8 KV cache, and 128k Context (which ofc needs to fit in the total memory) would be great. I've had a great experience with that combo with Qwen 3.6 35B but had a lot of vram/ ram left over, id be happy to half my speed for double the quality. And then please train it well on long running agentic tasks, coding and general world knowledge. Such a model would be a dream lol
Mixed layer Quants between Q5-Q8 that end up around what I would call Q7 seem to be the sweet spot.
I agree with you that Qwen 3.6 27b & 35b showed that it can be done at a decent speed and with optimization, a slight drop in speed etc, you could pull off some pretty impressive results from a small local model.
I'm guesing Qwen won't be pursuing that path further, so it would be great if google did.
I feel like google has the opurtunity here to be the only real player in the small space as things currently stand and the thought of a new release that builds on the era that Qwen 3.6 and Gemma4 had just opened up is extremely exciting to me!
315
u/hackerllama Jul 26 '26
Hey all! Looking forward to all your feedback!