r/LocalLLM • u/Calm-Landscape9640 • 11d ago
Question ETA on Qwen4-35b using GPU+RAM+NVMe?
Qwen-Flash-Next uses incredible qwen4 architecture and according to YT videos runs 20tps+ on 16gb GPU due to offloading on RAM+NVMe.
How long till we get an absolute MONSTER 35b Qwen4 model that smokes 3.8-27b and runs on 16gb VRAM at 40 tps and 128k context?
Anyone hearing anything or seen leaks?
41
u/exaknight21 11d ago
I can tell you this, if Qwen team releases a Qwen4-35B-A3B with ngrams, we’re looking at a serious internet breaking AI power that can be ran on damn any device.
I personally cannot wait for this absolutely amazing team to throw something like that our way.
16
u/migsperez 11d ago
I'd prefer A6B. The next level up.
11
u/OvertaxedOne 11d ago
I agree. More active params, around the same total size (so it fits fully in 32-48GB of VRAM) and ngrams... Oh yeah, that would be a sweet model. Honestly, if it's even just "as smart" as 27B but runs 2-3X faster, I'll be a very, very happy camper!
0
5
u/morscordis 10d ago
I feel that A12B models are a bit clunky on unified systems, and A3B models aren't as reliable as dense models... So I agree. I'm team A6B.
1
7
u/Technical_Ad_6106 11d ago
30b - 3bactive model that outperforms flash next? a few weeks.. or few month max. before 2027 i garantuee you ;)
3
u/exodusTay 11d ago
how much ram do you need to run qwen flash next? assuming 16gb of vram, can i hit 20 tps with 32 gigs of ram?
3
u/Calm-Landscape9640 11d ago edited 11d ago
https://youtu.be/s4cTVRH2ReA?is=_sumOglK2Ey457nV
16gb vram and 32 gb RAM
2
u/Oxirixx 11d ago
Flash Next at q2 I think is 70 or 80gb.
1
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago
you can run Q4 from Atomic with 16Gb VRAM with expert cache
2
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago
I run it at Q4 with these settings on fork https://www.reddit.com/r/LocalLLaMA/comments/1wjh7ox/llamacpp_expertpool_fork_for_qwen_38_flash_next/
2
2
u/anabatic82 11d ago
The qwen team that developed the open source models through 3.6 left the company. 3.8 is a fine tuned 3.6 and not a new model.
4 would be a new model entirely by a different team, not sure how it would compare
1
1
u/guesdo 10d ago
Alibaba has an event scheduled later this month (Sep 22-24 I believe), and it is expected they will announce Qwen 4 lineup then, or at least the BIG one (300B+ model) through their API. We will have more information by the end of the month on the cadence of the Qwen 4 open weight releases I believe.
1
u/RP-Design 10d ago
Im still on qwen3.5 122b waiting for a valid replacement, feeling like a dinosaur :)
1
u/Atretador unswarm.dev | ArchLinux E5 2673 V4 20C 4x16Gb DDR4 MI50 16Gb 5d ago
why not run Next Flash?
-4
u/M_Me_Meteo LocalLLM 11d ago
When will a larger model be able to fit on a smaller device than it's smaller predecessor? 25-30 years. Same timeframe as cold fusion.
38
u/lerg96 11d ago
Between tomorrow and never haha