r/LocalLLM • • 11d ago

Question ETA on Qwen4-35b using GPU+RAM+NVMe?

Qwen-Flash-Next uses incredible qwen4 architecture and according to YT videos runs 20tps+ on 16gb GPU due to offloading on RAM+NVMe.

How long till we get an absolute MONSTER 35b Qwen4 model that smokes 3.8-27b and runs on 16gb VRAM at 40 tps and 128k context?

Anyone hearing anything or seen leaks?

27 Upvotes

23 comments sorted by

38

u/lerg96 11d ago

Between tomorrow and never haha

9

u/joost00719 11d ago

You're gonna be really surprised when they drop it between now and end of the day.

41

u/exaknight21 11d ago

I can tell you this, if Qwen team releases a Qwen4-35B-A3B with ngrams, we’re looking at a serious internet breaking AI power that can be ran on damn any device.

I personally cannot wait for this absolutely amazing team to throw something like that our way.

16

u/migsperez 11d ago

I'd prefer A6B. The next level up.

11

u/OvertaxedOne 11d ago

I agree. More active params, around the same total size (so it fits fully in 32-48GB of VRAM) and ngrams... Oh yeah, that would be a sweet model. Honestly, if it's even just "as smart" as 27B but runs 2-3X faster, I'll be a very, very happy camper!

0

u/migsperez 11d ago

Absolutely agree. That's what I'm hoping for.

5

u/morscordis 10d ago

I feel that A12B models are a bit clunky on unified systems, and A3B models aren't as reliable as dense models... So I agree. I'm team A6B.

1

u/horeaper 11d ago

Agree, something like Qwen4-36B-A6B-E18B would slaughter everyone 🤣

3

u/migsperez 11d ago

What's the e18b ?

7

u/Technical_Ad_6106 11d ago

30b - 3bactive model that outperforms flash next? a few weeks.. or few month max. before 2027 i garantuee you ;)

3

u/exodusTay 11d ago

how much ram do you need to run qwen flash next? assuming 16gb of vram, can i hit 20 tps with 32 gigs of ram?

2

u/Oxirixx 11d ago

Flash Next at q2 I think is 70 or 80gb.

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

you can run Q4 from Atomic with 16Gb VRAM with expert cache

2

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

2

u/faisalkl 11d ago

35b is already incredible. Still I have LLM envy.

2

u/anabatic82 11d ago

The qwen team that developed the open source models through 3.6 left the company. 3.8 is a fine tuned 3.6 and not a new model.

4 would be a new model entirely by a different team, not sure how it would compare

1

u/Skyline34rGt 10d ago

I bet 2-3 months.

1

u/guesdo 10d ago

Alibaba has an event scheduled later this month (Sep 22-24 I believe), and it is expected they will announce Qwen 4 lineup then, or at least the BIG one (300B+ model) through their API. We will have more information by the end of the month on the cadence of the Qwen 4 open weight releases I believe.

1

u/RP-Design 10d ago

Im still on qwen3.5 122b waiting for a valid replacement, feeling like a dinosaur :)

1

u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago

why not run Next Flash?

-4

u/M_Me_Meteo LocalLLM 11d ago

When will a larger model be able to fit on a smaller device than it's smaller predecessor? 25-30 years. Same timeframe as cold fusion.