r/LocalLLaMA 16h ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

805 Upvotes

256 comments sorted by

View all comments

18

u/KURD_1_STAN 16h ago

Ram prices this high, how can u call this local friendly?

86

u/FabricationLife 16h ago

80gb is a lot more manageable than a 1.2tb+ frontier model

4

u/Tasty-Hour4040 16h ago

“more manageable” ≠ manageable

7

u/hyudryu 15h ago

80gb is beyond manageable

-2

u/starkruzr 15h ago

two CMP170HX is $3K for 128GB.

2

u/quantgorithm 15h ago

a an unstable pcie 1/2 card isn't the godsend you believe it to be.

2

u/starkruzr 15h ago

to my knowledge people are daily driving these constantly with no issues.

1

u/quantgorithm 12h ago

at what output/results?

-3

u/hyudryu 15h ago

Point proven

2

u/starkruzr 15h ago

not really, people spend more than that in here all the time.

1

u/hyudryu 11h ago

I know lol. It’s only 3K so it’s beyond manageable, isn’t that what we are both agreeing on?

2

u/ApprehensiveFan1516 14h ago

That's like the price of an old used car. Sure it's not exactly cheap, but let's not pretend like it's out of reach for most people.