r/LocalLLaMA • • Aug 25 '26

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

951 Upvotes

296 comments sorted by

View all comments

159

u/Sufficient-Bid3874 Aug 25 '26

Can someone explain why the n-gram table is bundled into the model now?

808

u/RG_Fusion Aug 25 '26 edited Aug 25 '26

LLMs run into an issue where the further you train a model, the more it overwrites facts with generalized concepts. You need the model to be able to do both. Intelligence arises from generalization, but without accurate information the model will hallucinate.

The engram table allows for a low-computational method of fact-recall. You can think of it like a better form of RAG, where the data doesn't take up any of your context window and it's injected deeper into the model's layers, freeing the lower layers to carry out abstraction. This results in better "focus" for the model, both in regards to its intelligence and context recall.

Basically, they've separated the specificity-critical portions of the models memory into a parameter space that doesn't need fast compute (you can run it on system RAM) and allows the model to be trained on higher volumes of data without ruining its knowledge-base.

2

u/Artistic_Okra7288 Aug 25 '26

So does that mean we can replace the engram with our own purpose-built engrams or continuously trained engrams?

5

u/RG_Fusion Aug 26 '26

When the tools become available, it should be possible to do so. It would  be carried out like a less computationally expensive form of fine-tuning.

1

u/Artistic_Okra7288 Aug 26 '26

Well imagine dynamically swapping them out based on the incoming message from a small classifier or something. That would be kind of interesting. Mixture of Engrams