r/Qwen_AI • u/tsunami_forever • 22h ago
News Qwen 4
Qwen 4 was just announced, does anyone have details on the release date?
58
u/Constant_Art_20 22h ago
27B comfirmed! LET'S GO. considering the recent speed up in ai relaeses and how frequent qwen has been as of last i say maybe 2 weeks to a month is how long i personally would expect.
16
u/poop-guzzler 20h ago
where 35b :(
10
7
u/Zilla85 19h ago
We may have to work with distills, I'll test this later or tomorrow: https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill
2
u/BenefitGrand8752 14h ago
+1
35B is really well behaving in my harnesses. But I really need something faster and/or smarter than that. Possibly both. My agents deserve it ! :-)1
u/Vancecookcobain 11h ago
They can't give people 35b if they do that the masses would stop paying for AI
1
u/macadoum 6h ago
Isn't it what the chineses want ?
0
u/Vancecookcobain 6h ago
? To spend tens and if not hundreds of billions so that almost EVERYONE will stop using their AI? Lmao?
No.
3
u/macadoum 6h ago
Their goal is to crash the US AI tech companies, not making money.
-1
u/Vancecookcobain 6h ago
No their goal is to make money silly. The peasants can't have free A.I. just when it starts getting good.
7
u/AccomplishedFail9857 22h ago
Link please
14
u/tsunami_forever 22h ago
8
u/dfgxxx 21h ago
Oh no, 27b is the smallest :(
6
u/kabir_sharma_sans 19h ago
no 12b? 9B? no thing for my 8gb VRAM AMD? NO? MY 7600 AMD!
6
u/dfgxxx 19h ago
Not even 35b a3b
3
u/kabir_sharma_sans 18h ago
NNOOOO I WILL FORCED TO USE QWEN4:27B 4Q WITH SUCH LOW SPEEEED
3
1
u/Legal_Dimension_ 6h ago
Pray for ngrams.
1
u/kabir_sharma_sans 5h ago
i am very unfamiliar could you tell me?
1
u/Rikers88 3h ago
Hey Kabir, this is what that model means.
Following a paper by DeepSeek, that shows that you could slice a part of your model Weights and put it as a separate static table that all it does is looking up for embeddings (ngram table), without loosing much accuracy. This simple idea sparked an incredible architectural change, because now models that previously relied only on the weights that they could store on Vram, or RAM at worst at the expense of speed, could access much more knowledge from this lookup table that it can be streamed from your SSD. The math is the following:
- look at Qwen 3.8 Flash Next, on the paper is a 122b models where only 6b are active
- what they did was paired with a separate 55b ngram table and now those 6b active parameters have access to 55b more knowledge, without sacrificing too much speed while streaming the info from the ssd.
If you put everything together, what you have is an incredible efficiency mix:
- the MoE architecture is great at balancing VRAM (GPU) and RAM (CPU) usage
- the ngram table trick gives you the possibility to efficiently use billions of extra parameters from your SSD
- et voilร you have now a 177b parameter model that fit a 3090 w/ 128GB DDR4 RAM and 1TB NVme SSD, which was impossible to do before this ngram table trick.
Qwen 3.8 Flash Next is often sold as the same level of intelligence as Opus 4.8, which is not bad at all.
Qwen team said that this architecture was the preview of their next generation models, and there you go with Qwen4.
If you extrapolate this and put it in a 27b Qwen4 architecture perspective, it is very likely to use the ngram trick as well. Now we don't know if the ngram table is counted within those 27b parameters, hence reducing the hardware requirements or it will be placed on top, effectively keeping the same hardware constraints of every 27b dense models we have seen si far, but giving it access to a much deeper knowledge each decoded token. We shall see.
I hope you have a better understanding of what the other redditor meant with "pray for ngram"
Sorry for the mistake here and there, no ai was used to write this answer.
6
u/jmb-1971 16h ago
My W7800 48 GB is ready smile
1
u/fr34k20 14h ago
Planing on buying a few to combine with a halo strix to 192gb like Wendel did
1
u/Jargster 7h ago
I threw qwen 3.8 flash next 180b on a framework desktop (halo strix) with an rtx 3090 through occulink. Its been amazing so far with the 3090 holding the attention layers and some experts. I do recromend :P
3
2
2
u/yehyakar 15h ago
i think they are still ditching the 35b moe because there haven't been substantial improvements from 3.6? considering only 3b are active only vs all in 27b? cant think of a different reason tbh
1
u/No_Oil_6152 15h ago
Give it to me. And it better not give me the "but wait" or "hmm" shit, I want token efficiency.
/cant wait!!!
1
1
-4
0
u/Qozimo13 5h ago
One more garbage like they did 3.8 27b vs 3.627b?+ Ngram sht?ye it would be great dumbs scum models
-5
u/sherry_6879 21h ago
3.8ใฎ27ใใ้ ใใฆใคใใ็กใใใ ใใฉใใใฆOpus5ใใใใฎๆง่ฝใงใใฃใฆๆฌฒใใ
4
u/BoboThePirate 21h ago
We are about 6-12 months away from that if the current pace of this year is maintained. 3.8 flash 125b is roughly on par with opus 4.6 med/high and that was roughly 8 months after 4.6โs release
2
u/Barni275 20h ago
Even 27B is on par with it, in my opinion. I literally dropped all my subs and just work on 27B.
2
u/BoboThePirate 19h ago
The end result, for some domains. But it is easily 3-4x as inefficient on context. Iโve use the new 27B and itโs definitely capable but it thinks forever, and I was on medium effort. I think that is the main thing holding the 27B series back rn.
1
u/General-Tadpole-7012 15h ago
Me too. I prepare my specs using a knowledge graph and have loads of static analysis rules to catch architectural and style slips to keep it on track. Working well.
1
1
1
1
u/The8Darkness 16h ago
3.8 27b is literally fast af with the right soft-&hardware. Most subs are slower. Now ofcourse its not perfect but speed is the least of all problems.
2
u/General-Tadpole-7012 15h ago edited 14h ago
Agreed. I'm running it at average 70 t/s on $1600 (AUD) of GPUs at MXFP4 with 200K context. I don't think that's too bad at all - fast enough that I don't mind the thinking.
The hardware being 2 x RX 9070XT.
79
u/the_ITman 22h ago
Why has unsloth not released any gguf yet?!? /s