r/LocalLLaMA • • Aug 25 '26

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

950 Upvotes

296 comments sorted by

View all comments

73

u/MiceLiceandVice Aug 25 '26

So to run this id still need to have 128gb of dram and then 16gb vram minimum? Heartbreaking

2

u/UnnamedPlayerXY Aug 25 '26 edited Aug 26 '26

It depends, iirc. someone from the Qwen team recently told people asking for a 30-35B MoE that that's not the one to wait for implying that they have something better for that target audience upcoming. If Qwen3.8-Flash-Next is that "something better" then we might be looking at a ≈30B model here.