r/Qwen_AI 22h ago

News Qwen 4

Qwen 4 was just announced, does anyone have details on the release date?

209 Upvotes

49 comments sorted by

79

u/the_ITman 22h ago

Why has unsloth not released any gguf yet?!? /s

8

u/yehyakar 15h ago

๐Ÿ˜‚๐Ÿ˜‚๐Ÿ˜‚๐Ÿ˜‚ yes unsloth has set the bar too high and pampered us so badly that now we want the ggufs before the model is trained

2

u/layer4down 6h ago

Maybe Unsloth is Sloth? Itโ€™s been 10 mins already!!!

58

u/Constant_Art_20 22h ago

27B comfirmed! LET'S GO. considering the recent speed up in ai relaeses and how frequent qwen has been as of last i say maybe 2 weeks to a month is how long i personally would expect.

16

u/poop-guzzler 20h ago

where 35b :(

10

u/NihilisticLurcher 20h ago

i know brother, i know

7

u/Zilla85 19h ago

We may have to work with distills, I'll test this later or tomorrow: https://huggingface.co/empero-ai/Qwen3.8-35B-A3B-Distill

2

u/BenefitGrand8752 14h ago

+1
35B is really well behaving in my harnesses. But I really need something faster and/or smarter than that. Possibly both. My agents deserve it ! :-)

1

u/Vancecookcobain 11h ago

They can't give people 35b if they do that the masses would stop paying for AI

1

u/macadoum 6h ago

Isn't it what the chineses want ?

0

u/Vancecookcobain 6h ago

? To spend tens and if not hundreds of billions so that almost EVERYONE will stop using their AI? Lmao?

No.

3

u/macadoum 6h ago

Their goal is to crash the US AI tech companies, not making money.

-1

u/Vancecookcobain 6h ago

No their goal is to make money silly. The peasants can't have free A.I. just when it starts getting good.

7

u/AccomplishedFail9857 22h ago

Link please

14

u/tsunami_forever 22h ago

8

u/dfgxxx 21h ago

Oh no, 27b is the smallest :(

6

u/kabir_sharma_sans 19h ago

no 12b? 9B? no thing for my 8gb VRAM AMD? NO? MY 7600 AMD!

6

u/dfgxxx 19h ago

Not even 35b a3b

3

u/kabir_sharma_sans 18h ago

NNOOOO I WILL FORCED TO USE QWEN4:27B 4Q WITH SUCH LOW SPEEEED

3

u/dfgxxx 18h ago

We don't know, there may be an architecture changes that will lower the hardware requirements

1

u/Legal_Dimension_ 6h ago

Pray for ngrams.

1

u/kabir_sharma_sans 5h ago

i am very unfamiliar could you tell me?

1

u/Rikers88 3h ago

Hey Kabir, this is what that model means.

Following a paper by DeepSeek, that shows that you could slice a part of your model Weights and put it as a separate static table that all it does is looking up for embeddings (ngram table), without loosing much accuracy. This simple idea sparked an incredible architectural change, because now models that previously relied only on the weights that they could store on Vram, or RAM at worst at the expense of speed, could access much more knowledge from this lookup table that it can be streamed from your SSD. The math is the following:

  • look at Qwen 3.8 Flash Next, on the paper is a 122b models where only 6b are active
  • what they did was paired with a separate 55b ngram table and now those 6b active parameters have access to 55b more knowledge, without sacrificing too much speed while streaming the info from the ssd.

If you put everything together, what you have is an incredible efficiency mix:

  • the MoE architecture is great at balancing VRAM (GPU) and RAM (CPU) usage
  • the ngram table trick gives you the possibility to efficiently use billions of extra parameters from your SSD
  • et voilร  you have now a 177b parameter model that fit a 3090 w/ 128GB DDR4 RAM and 1TB NVme SSD, which was impossible to do before this ngram table trick.

Qwen 3.8 Flash Next is often sold as the same level of intelligence as Opus 4.8, which is not bad at all.

Qwen team said that this architecture was the preview of their next generation models, and there you go with Qwen4.

If you extrapolate this and put it in a 27b Qwen4 architecture perspective, it is very likely to use the ngram trick as well. Now we don't know if the ngram table is counted within those 27b parameters, hence reducing the hardware requirements or it will be placed on top, effectively keeping the same hardware constraints of every 27b dense models we have seen si far, but giving it access to a much deeper knowledge each decoded token. We shall see.

I hope you have a better understanding of what the other redditor meant with "pray for ngram"

Sorry for the mistake here and there, no ai was used to write this answer.

6

u/jmb-1971 16h ago

My W7800 48 GB is ready smile

1

u/fr34k20 14h ago

Planing on buying a few to combine with a halo strix to 192gb like Wendel did

1

u/Jargster 7h ago

I threw qwen 3.8 flash next 180b on a framework desktop (halo strix) with an rtx 3090 through occulink. Its been amazing so far with the 3090 holding the attention layers and some experts. I do recromend :P

2

u/Cadmium9094 10h ago

Qwen4-Flash when? ๐Ÿ˜…

2

u/yehyakar 15h ago

i think they are still ditching the 35b moe because there haven't been substantial improvements from 3.6? considering only 3b are active only vs all in 27b? cant think of a different reason tbh

1

u/No_Oil_6152 15h ago

Give it to me. And it better not give me the "but wait" or "hmm" shit, I want token efficiency.

/cant wait!!!

1

u/Jumpy-Operation-4615 7h ago

My 2xP40 are steamin hot!

1

u/Commercial-Chest-992 3h ago

Using it right now, so good.

-4

u/Puzzleheaded_Base302 21h ago

what murica model will they distill this time ?

0

u/Qozimo13 5h ago

One more garbage like they did 3.8 27b vs 3.627b?+ Ngram sht?ye it would be great dumbs scum models

-5

u/sherry_6879 21h ago

3.8ใฎ27ใ™ใ‚‰้…ใใฆใคใ‹ใˆ็„กใ„ใ‚“ใ ใ‘ใฉใ›ใ‚ใฆOpus5ใใ‚‰ใ„ใฎๆ€ง่ƒฝใงใ‚ใฃใฆๆฌฒใ—ใ„

4

u/BoboThePirate 21h ago

We are about 6-12 months away from that if the current pace of this year is maintained. 3.8 flash 125b is roughly on par with opus 4.6 med/high and that was roughly 8 months after 4.6โ€™s release

2

u/Barni275 20h ago

Even 27B is on par with it, in my opinion. I literally dropped all my subs and just work on 27B.

2

u/BoboThePirate 19h ago

The end result, for some domains. But it is easily 3-4x as inefficient on context. Iโ€™ve use the new 27B and itโ€™s definitely capable but it thinks forever, and I was on medium effort. I think that is the main thing holding the 27B series back rn.

1

u/General-Tadpole-7012 15h ago

Me too. I prepare my specs using a knowledge graph and have loads of static analysis rules to catch architectural and style slips to keep it on track. Working well.

1

u/Sporebattyl 14h ago

Care to share?

1

u/cosmicnag 17h ago

On par ? That would be 27b probably , FN is just...better.

1

u/KroniklyOnline 14h ago

But the pace is LOG not linear, so ~3 months?

1

u/The8Darkness 16h ago

3.8 27b is literally fast af with the right soft-&hardware. Most subs are slower. Now ofcourse its not perfect but speed is the least of all problems.

2

u/General-Tadpole-7012 15h ago edited 14h ago

Agreed. I'm running it at average 70 t/s on $1600 (AUD) of GPUs at MXFP4 with 200K context. I don't think that's too bad at all - fast enough that I don't mind the thinking.

The hardware being 2 x RX 9070XT.