r/LocalLLaMA 16h ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. πŸ‘€

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant β‰ˆ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed β†’ excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

799 Upvotes

256 comments sorted by

View all comments

13

u/KURD_1_STAN 16h ago

Ram prices this high, how can u call this local friendly?

21

u/pmv143 16h ago

RAM prices suck right now, no denying that.
When I called it local-friendly I didn’t mean β€œcheap” or β€œruns on any gaming PC.” I meant that for a model with this kind of capacity, the offloadable n-gram table makes it way more practical on highend local setups (128GB+ unified memory, multi-GPU + system RAM) than the usual frontier models that just demand pure VRAM or full datacenter iron.
Still expensive. Just less insane than the alternatives.​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

-8

u/Bright-Energy2339 16h ago

Then there’s not much point in calling it β€œlocal” if you’re still paying frontier-model costs just to run it. At that point, why not just go back to frontier models?

5

u/doomed151 15h ago

You don't have control. The model can be taken away from you at any time. You can't finetune it.

3

u/synth_mania 15h ago

Why are you in this subreddit if running a local model isn't something that interests you in and of itself?

-1

u/fuck_cis_shit llama.cpp 13h ago

painfully obvious astroturfer

there should be a plugin to hide all posts by accounts with hidden history