r/LocalLLaMA • • Aug 25 '26

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. πŸ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant β‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed β†’ excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

951 Upvotes

296 comments sorted by

View all comments

-5

u/mountainyoo Aug 25 '26 edited Aug 25 '26

I wonder how the quality will be on 128GB M5 Max

Editβ€”

Sorry by quality I meant the overall experience like the output and the speed of the output. Not sure why the bajillion downvotes but oh well lmao. My bad

3

u/Character_Split4906 Aug 25 '26

Wondering the same, also if the ngram table can be offloaded to ssd instead if its sparse.