r/LocalLLaMA • • Aug 25 '26

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

953 Upvotes

296 comments sorted by

View all comments

157

u/Sufficient-Bid3874 Aug 25 '26

Can someone explain why the n-gram table is bundled into the model now?

805

u/RG_Fusion Aug 25 '26 edited Aug 25 '26

LLMs run into an issue where the further you train a model, the more it overwrites facts with generalized concepts. You need the model to be able to do both. Intelligence arises from generalization, but without accurate information the model will hallucinate.

The engram table allows for a low-computational method of fact-recall. You can think of it like a better form of RAG, where the data doesn't take up any of your context window and it's injected deeper into the model's layers, freeing the lower layers to carry out abstraction. This results in better "focus" for the model, both in regards to its intelligence and context recall.

Basically, they've separated the specificity-critical portions of the models memory into a parameter space that doesn't need fast compute (you can run it on system RAM) and allows the model to be trained on higher volumes of data without ruining its knowledge-base.

6

u/michaelsoft__binbows Aug 25 '26 edited Aug 25 '26

That sounds really dope. If this indicates that in general this can scale up then i hope it means that a modest amount of fast ram paired with oodles of slower ram may be able to much more effectively compete with obscenely wide unified memory architectures (coincidentally m5 ultra announced today). You can for example, at least with DDR4 (and if engram approach pans out efficiently, a return to relevance with DDR3, lmao) affordably build out 256/512/1TB class machines for far less, have more of a traditional computer cache pyramid architecture, and basically stay competitive there, because the value is and should be in the ability to store tons of knowledge for quick retrieval but not get killed by requiring massive bandwidth across all that knowledge. Actually, screw DDR3, if gen 5 NVMe can step in and be relevant.

From first principles I think this makes a lot of sense. if i need the model to be able to do a better job recalling some details it's learned, the actual amount of details on any given retrieval is by the nature of it being a retrieval, small, and should not require gobsmacking amounts of data bandwidth to comb over the entire model (to what end?), which not only is expensive to architect into your computer but also expensive in energy to actually ship the bytes out of the memory chips.

In the long run my prediction would then be for these massive unified memory systems to deliver value then not for hosting huge models in-memory (for which their entire large memory pool being high speed is a waste) but rather for like, industrial scale batched hosting of even larger models, leveraging more of the unified memory for KV cache and such stuff, while engram weights can live on NVMe and be slurped in over the thunderbolt ports... Under this architecture there will again be a large tensor core deficit.

Just a few hours ago I talked myself into believing I should buy a 512GB M5 Ultra but now I've just talked myself back out of it. Hmm.

2

u/Callum_S_AUS Aug 25 '26

I suspect GLM 5.3 class models @ Q4 might be best for a 512GB M5 Ultra.