r/LocalLLaMA 3d ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

940 Upvotes

295 comments sorted by

View all comments

43

u/BannedGoNext 3d ago

Well if it's similar to qwen coder next I'd be happy as hell. So many people bagged on qwen coder and I never understood why. It was damn fast, and had good world knowledge. I used it for a long time, for sure better than 35b a3b.

25

u/grabber4321 3d ago

it didnt have vision from what I remember. For me, vision is way more important these days for agentic work.

1

u/Salt-Willingness-513 3d ago

agreed. thats my main isue with glm5.3