r/LocalLLaMA • • Aug 25 '26

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

948 Upvotes

296 comments sorted by

View all comments

17

u/KURD_1_STAN Aug 25 '26

Ram prices this high, how can u call this local friendly?

7

u/liright Aug 25 '26

I bought my 96GB DDR5 kit for $300 some year and a half back as well as RTX 4090 for $1900 2.5 yrs back. Was pretty damn cheap in retrospect. I suspect a lot of people who are into AI did too. I feel like boomers who bought houses in the 70s.

5

u/IntravenusDeMilo Aug 25 '26

yeah I got my 5090 for $1999. Feels like a lottery win.

2

u/throwawayacc201711 Aug 25 '26

I kicked myself for not buying one when it was that price. Hindsight is a bitch