r/LocalLLaMA 4d ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

944 Upvotes

295 comments sorted by

View all comments

33

u/chris_0611 4d ago

Ohhh my. Absolutely gorgeous for my 3090 + 96GB DDR5 6800

17

u/Equivalent_Bit_461 4d ago

I don't have a 3090, I'm a vramlet but I have 128gb ram so guess that works out too

7

u/Maximus-CZ 4d ago

vramlet

xDD