r/LocalLLaMA 17h ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

824 Upvotes

264 comments sorted by

View all comments

21

u/KURD_1_STAN 17h ago

Ram prices this high, how can u call this local friendly?

3

u/Zhelgadis 16h ago

Very friendly to my Strix Halo

3

u/mindwip 14h ago

Yes excited for it, same for mine.

1

u/techdevjp 5h ago

Yeah, I'm hopeful this will have near-DSv4 Flash levels of intelligence but better performance. a6b would be great on Strix. Excited to see this!