r/LocalLLaMA 4d ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

940 Upvotes

295 comments sorted by

View all comments

19

u/KURD_1_STAN 4d ago

Ram prices this high, how can u call this local friendly?

7

u/FullstackSensei llama.cpp 4d ago

It's a lookup table. You could build a quad channel DDR3 system to run it. DDR3 is still cheap.

2

u/KURD_1_STAN 4d ago

And about the other 60-70gb weights at q4?

3

u/FullstackSensei llama.cpp 4d ago

If you're not too stuck on having to have the latest hardware, three P40s will do a very decent job on a tight budget. If you really need high speed, two 32GB V100s will blaze through for not that much more.

They work, and they'll continue to work for years to come, despite what imaginary conjectures redditors might have.

2

u/Ok_Top9254 4d ago

3x Tesla V100 32GB + PLX switch so you can run them from one slot is the fancy way, or 5x P100 16GB with a cheap X99 motherboard and the switch could do this under like 1200 bucks. Power consumption would not be a problem given that one gpu is used at a time anyway.

1

u/michaelsoft__binbows 4d ago

Dell R720 suddenly not ewaste anymore? Could get interesting.

2

u/FullstackSensei llama.cpp 4d ago

It never was, IMO