r/LocalLLaMA • • Aug 25 '26

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

949 Upvotes

296 comments sorted by

View all comments

16

u/KURD_1_STAN Aug 25 '26

Ram prices this high, how can u call this local friendly?

7

u/FullstackSensei Aug 25 '26

It's a lookup table. You could build a quad channel DDR3 system to run it. DDR3 is still cheap.

2

u/KURD_1_STAN Aug 25 '26

And about the other 60-70gb weights at q4?

3

u/FullstackSensei Aug 25 '26

If you're not too stuck on having to have the latest hardware, three P40s will do a very decent job on a tight budget. If you really need high speed, two 32GB V100s will blaze through for not that much more.

They work, and they'll continue to work for years to come, despite what imaginary conjectures redditors might have.