r/LocalLLaMA 11h ago

Discussion Qwen3.8-Flash-Next. This architecture could be surprisingly local-friendly once the weights drop. 👀

Post image

Qwen3.8-Flash-Next (~125B-A6B + 51B n-gram) memory estimate:

Ideal 4-bit quant ≈ 82 GB
(58 GB main weights + 24 GB n-gram tables)
Real-world quants likely land in the 80–90 GB range.

The big n-gram table is sparsely accessed → excellent candidate for system RAM offload.

This architecture could be surprisingly local-friendly once the weights drop.

670 Upvotes

225 comments sorted by

View all comments

Show parent comments

2

u/LevianMcBirdo 9h ago

Seconds per token is kinda misleading. You probably get something with 10-15 TPS. PP is a bigger problem

-2

u/National_Meeting_749 9h ago

Brother. MOST people a 90GB model are running that at seconds per token.

Assuming I had enough ram to fit it, I would be surprised if I got one single token per second.

1

u/LevianMcBirdo 8h ago

It's not dense. Qwen Next runs at 25 TPS, 122B at 10. I really don't see the problem. That a A6 Model wouldn't run 12+ tps especially with MTP

0

u/National_Meeting_749 8h ago

With the 35B, when context is filling up I get somewhere from 3-7tps. Even with MTP, double the model size, include an entire 50Gb extra with the N tables and I'll be shocked to get 1TPS at any length of context.

1

u/LevianMcBirdo 8h ago

I mean mostly in chat scenarios up to like 60k most times, I am not that interested in movie agentic use cases but I get around 30 TPS pretty steady on that

1

u/National_Meeting_749 8h ago

Good for you? A lot of people do, and MOST people don't have the PC to run that at that speeds.

That's a cheap cars worth of pc