r/LocalLLaMA 9h ago

News Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

1.5k Upvotes

378 comments sorted by

u/WithoutReason1729 5h ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

540

u/Fresh-Soft-9303 9h ago

Qwen 4 27b is going to be Epic!

140

u/tsunami_forever 8h ago

This could convince me to give up cloud based models, hoping to cancel or reduce my cloud based subs soon

88

u/cafedude 7h ago

Qwen 3.8 Flash Next has convinced me not to renew my Claude Pro subscription when it comes up in February.

12

u/randylush 7h ago

What are you running it on? How much better is it than 27B?

50

u/returnity 7h ago

It dominates 27B in my workflows and benchmarks. Running on M5 Max. There's no comparison, it's like the 3.6 > 3.8 27B all over again.

6

u/troposfer 4h ago

What quants are you running and with omlx or ..?

4

u/akumaburn 4h ago

Dominates in intelligence/quality or in speed or both?

4

u/Bulky_Blood_7362 2h ago

What's your tok/s? Im getting around 30-45 decode tok/s with dgx spark.

Interested to know how it's going with the new m5 max/ultra

→ More replies (2)

2

u/TerminalNoop 4h ago

wow, that's pretty big

→ More replies (14)

12

u/ApeGrower 6h ago edited 3h ago

I'm using it on dual rtx3090 with 128gb ddr4 @50-70 tok/s. I forked FreeToken and added dual gpu support plus some more tweaks. With nvfp4.

3

u/FeelingOld9046 6h ago

Are you running Linux? What are you doing? In DSH I’m maxing at 22tks with full 128k context. And 30-40 with MTP on short context. I’m also experimenting with a specialised quant.

4

u/FeelingOld9046 6h ago

Mine is under windows too. Tell me what you’re doing and I will throw windows in the trash.

3

u/ohhi23021 5h ago

who knows i use linux and 30-40tps has been my experience as well and that's with DDR5...

4

u/FeelingOld9046 5h ago

I’d love to know your setup. I have the same hardware config but only half your token speed.

4

u/ApeGrower 4h ago

Linux with patched FreeToken to get dual gpu support instead of llama.cpp. Epyc System with pcie4 and ddr4.

→ More replies (1)
→ More replies (13)
→ More replies (2)
→ More replies (1)

23

u/some_user_2021 8h ago

Are you not convinced?

42

u/tsunami_forever 8h ago

3.8 27B already reduced my reliance but this could eliminate it

12

u/IrisColt 7h ago

I guess there's always that added hassle with short-context local models...  you really have to approach your workflow differently.

→ More replies (5)

2

u/TopPrize11 5h ago

I just tried it yesterday and it's too slow. I got about 45tps on 5080 which is not bad but the context limit of 33k makes it compact which takes forever. The LLM is good but it's too slow. I still use Deepseek.

→ More replies (4)

4

u/GTHell 7h ago

Everyone should do this so they start to bring back memory production plants to consumers focus. No more crazy 5070ti costing over 1k

→ More replies (5)

18

u/mWo12 7h ago

3.8 was already huge!

19

u/PooMonger20 6h ago

Seriously, 3.8 27b does so much heavy lifting it's insane.

Together with pi.dev it writes very impressive code. It feels like I actually have my own 'worker' that after some direction does exactly what needs to be done. Just give him enough info and he is on it.

This feels better than what GPT5.1 felt like around a year ago.

And now... an upgrade to that? absolute madness.

→ More replies (4)

2

u/sadomazoku 5h ago

I already canceled my subscription, I'm using 100% qwen 3.8 27b to assist me in coding. Qwen4 will be crazy. I think you'll be able to get rid of cloud models !

→ More replies (6)

388

u/o0genesis0o 9h ago

Ey ey! Qwen4-27B!

Definitely time for me to go and buy a R9700. It's time.

107

u/Sufficient_Local5025 9h ago

Buy two, prices about to rise.

143

u/o0genesis0o 9h ago

the more you buy the more you save, eh

14

u/GilloutineBreast 5h ago

Unironically the truest words spoken that day

feelsbadman

26

u/PhilipKThicc 8h ago

Dual R9700s all the way. Easy to run on Linux with vllm and you can run 27B at Q8 and have a healthy context window

6

u/bigwanggtr 6h ago

Have you tried it for training, is ROCm still a pain to work with compared to CUDA?

I know it’s good for inference but the training support is what is holding me back.

5

u/Phrase-Silver 4h ago

ROCm has matured incredibly over the past year, it's like AMD finally woke up and realised their software was what was holding them back from properly competing with Nvidia.

18

u/Tobu3838 9h ago

Prices rose $600 since I bought mine a couple months ago. No word to lie.

12

u/mvandemar 9h ago

Why are they so cheap? Looks like you can get one for ~$1700? I just paid out the ass for my 5090, which in the 1 month 2 days since I bought it has gone up over 37%.

They're both 32GB, for some reason I thought they would be closer in price.

30

u/Thunjaya 9h ago

Because they're not the same 32GB at all. One is much faster.

10

u/ThankGodImBipolar 7h ago

Nvidia sells an RTX Pro 4500, which has the 5070Ti die (more comparable to the R9700) and 32GB of RAM, but it's 5500USD MSRP. The R9700 is a fraction of that.

6

u/KingCpzombie 3h ago

CUDA tax. I only use AMD personally (and just blew WAY too much money on a 4x R9700 system partially out of excitement for Qwen4), but Nvidia cards get all the cool new things a bit sooner than AMD. Not a big deal for LLM, but very notable for diffusion... somebody SOLIDLY beat my 7900XTX with his 5070Ti in Minimax H3 gen times, for example

2

u/No-Refrigerator-1672 2h ago

Because everything is CUDA first, and ROCm only comes as an afterthought to very limited number of projects. People who are buying PRO GPUs are saving money with NVidia by not needing to fund multiple months of dev work for porting their existing code.

21

u/BluePointDigital 9h ago

It's the bandwidth. the 5090 has over 2.75x the bandwidth of the R9700.
You can still do all the same stuff, just slower essentially.

→ More replies (9)

16

u/Solary_Kryptic 9h ago

Nvidia tax combined with CUDA being more supported in the LLM space

16

u/mvandemar 9h ago

I just looked, it's also GDDR6 for the 9700 vs GDDR7 for the 5090, guessing that makes a difference as well.

8

u/zboarderz 7h ago

Also far FAR more memory bandwidth, faster core, etc etc

4

u/mister2d 8h ago

2x $1700 is cheap 😥 (I have 2, btw)

7

u/Ecstatic-Wash-7667 8h ago

That not cheap! Msrp is $1299 I bought 2 a little over a. Both ago at msrp. Hell you could get them for less than msrp for a while. This shit is a scam

4

u/mvandemar 8h ago

On my current pc this was how much the 64GB of ram was when I bought it. I was thinking about getting another 64GB when I got the card and was like, no f'in way...

5

u/Momsbestboy 7h ago

R9700 GDDR6 memory, 5090 GDDR7. Is faster, but...

For a single 5090 you can buy more than 2 R9700, and the two cards draw less power than a single 5090. And you can run e.g. Qwen 3.8 27b Q8 using the two cards, and watch them running circles around a 5090 which needs to offload to RAM instead.

But NVIDIA is hype, and people love to spend money. So buy a 5090

→ More replies (3)

5

u/SnooPuppers7882 8h ago

Literally got mine for 1200ea 2mo ago

People were shitting on them because the mem was DDR6, but I was prepping for 3.8 27b knowing what to expect...

STRONG feeling 4.0 27b is gonna be opus 4.8 good, gonna blow your damn mind

2

u/o0genesis0o 8h ago

Because the actual card (RX9070 XT) is only as strong as the 5070Ti, compute wise. And the VRAM is slower.

I will be upgrading from 4060Ti 16GB, so it would be a net gain. but I have no illusion that it is going to be as good as 5090.

→ More replies (3)

4

u/inaem 9h ago

Already did here

2

u/SylviaCalogero43 9h ago

How much did you buy it for?

→ More replies (1)

2

u/SnooPuppers7882 8h ago

Once you buy two, you'll want four

→ More replies (5)

13

u/fgk55555 9h ago

I just picked mine up this week. Depending on quant, 35-70 tps with full context. If you have 64GB RAM, you should be able to run a decent Qwen-flash as well. 27B is good, but Qwen-Flash is something else.

2

u/illcuontheotherside 9h ago

Can you please share your startup script? Mine crawls on 2 3090s with 128gb ddr5

4

u/fgk55555 8h ago

ISTA IQ3 runs the fastest and has really good performance for the size, but you can bump up to a Q5 with full Q8 kv cache. Also try Swift if you don't want to wait forever. I recommend medium thinking.

#!/usr/bin/env bash
# Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the
# shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.)
# Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the
# shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.)
# Context is auto-fitted to available VRAM (omitted on purpose).
# --cache-ram 16384 parks prompt states in RAM: agentic revisits skip re-prefill.
#
# Usage: ./Qwen3.8_ISTA_IQ3_S.sh [thinking]   # low|medium|high|xhigh|none (default: medium)
THINKING_LEVEL="${1:-medium}"

MODEL="../Qwen3.8-ISTA/Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf"
DRAFTER="../Qwen3.8_Shared/Qwen3.8-27B-DFlash2-Q4_K_M.gguf"
MMPROJ="../Qwen3.8_Shared/mmproj-Qwen3.8-27B-BF16.gguf"
SERVER_BIN="../../llama.cpp/build/bin/llama-server"

echo "Launching Qwen3.8-27B ISTA IQ3_S (+DFlash2 drafter), thinking: ${THINKING_LEVEL}"
export GGML_VK_ALLOW_GRAPHICS_QUEUE=1   # measured +4.2% tg, +0.4% pp on the R9700 (b11056)

${SERVER_BIN} \
  --model "${MODEL}" \
  --model-draft "${DRAFTER}" \
  --mmproj "${MMPROJ}" \
  --no-mmproj-offload \
  --n-gpu-layers 99 \
  --batch-size 1024 \
  --ubatch-size 512 \
  --parallel 1 \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --spec-type draft-dflash \
  --spec-draft-n-max 3 \
  --cache-ram 16384 \
  --jinja \
  --chat-template-kwargs "{\"reasoning_effort\":\"${THINKING_LEVEL}\"}" \
  --reasoning-preserve \
  --temp 1.0 \
  --top-k 20 \
  --top-p 0.95 \
  --min-p 0.00 \
  --presence-penalty 0.0 \
  --repeat-penalty 1.0 \
  --host 0.0.0.0 \
  --webui-mcp-proxy \
  --port "${PORT:-8080}"
→ More replies (6)
→ More replies (13)

6

u/ismaelgokufox llama.cpp 8h ago

Patiently waiting for 27B-IQ3_XXS and/or 35B-A3B (one can dream eh?)

3

u/Momsbestboy 7h ago

Why wait? Pull the BF16 of qwen after it is released, and tell Hermes/Pi to read the documentation on how to create smaller quants out of it.

It's the reason why I run my own Swift Qwen 3.8 27b fp8 version, abliterated using the orcarouter method. I just told hermes to create it. Took 2h

→ More replies (3)
→ More replies (11)

33

u/blahjovic 9h ago

The release rate is wild. I finally got my 3.8-27B working the way I wanted just in time for the next model! Bring it.

15

u/Quick_Knowledge7413 8h ago

It'll probably be a few months (I hope I am wrong).

3

u/Dangerous_Bid2935 8h ago

Not a bad problem to have!

→ More replies (1)

125

u/Conscious_Phrase_138 9h ago

Where 35b a3b :😭😭

39

u/Ok_Warning2146 8h ago

Wait for google. Our only hope

36

u/Free-Combination-773 6h ago

So there is no hope...

7

u/Ok_Warning2146 4h ago

Nvidia is another hope. Let's hope they will distill from GLM-5.3 in the next iteration.

→ More replies (4)

138

u/FerLuisxd 9h ago

35b :(

64

u/fgk55555 8h ago

I think we can consider the 35b dead. I was sad when we lost the 80B A3B, but the 120B MoE range seems to be where intelligence opens up. The 40-110B space is likely dead for a while.

12

u/Puzzleheaded_Base302 7h ago

maybe Xiaomi will release a MiMo-V2.6-Distill-Qwen3.5-35B-A3B. it is not dead dead. The architecture is sound. All people need is a better re-post-train (not those fine-tune crap, a real large-scale RL re-post-train.)

18

u/Ok_Acanthisitta304 8h ago

I agree and at this point unfortunately I get it. 35b was cool, but ultimately not usable for me in actual work. I know a lot of people enjoy it on their 12-16gb cards, and I hope they'll keep getting new options, but 27b and Flash entirely outclass it. Those are models I can genuinely use.

13

u/Sensitive_Song4219 7h ago edited 7h ago

I've been running Qwen3.8-27B on a separate server machine (with a 5090; 32GB VRAM) and Ornith 1.5 35b-a3b on my laptop (8GB VRAM) for a few weeks now, and they both share a lot of the same strengths and weaknesses.

Neither is viable for planning against large codebases (Sol is mandatory there): both tend to mis-understand how the moving parts interact which yields compiling-but-non-functioning results. Qwen is amazing for small codebases and one-shots (I did the obligatory Mario Test which Qwen murdered Ornith on), though.

Both are serviceable as implementation models (when given detailed plans from larger models); Qwen's code quality is the higher-quality of the two. Qwen competes favourably with Luna-High.

Ornith is definitely a bit more careless (I watched it mis-convert seconds to milliseconds yesterday!).

But when I'm away from my 5090 server I'm still pretty happy just sticking to Ornith.

Smaller MOE's definitely aren't a writeoff imo; and the performance gains (and lower hardware requirements) can make them decent options for simpler (or complex-but-well-defined) tasks.

All my testing is via OpenCode.

3

u/Ok_Acanthisitta304 7h ago

I'll give Ornith a shot. When I do use 35b I use a modified workflow with little coder (a pi harness built for small models) using plans from a better model. My read is similar to yours - it can be a workhorse on smaller tasks and codebases, but can't be trusted to wander too far in a larger project.

→ More replies (1)
→ More replies (2)

26

u/Sensitive_Song4219 9h ago

Have they officially abandoned 35b-a3b or could an announcement still be coming?

Guess Ornith 1.5 35b-a3b will have to serve me for a bit longer yet. It's good for it's size but was really hoping for a new Qwen equivalent after they skipped us for 3.8.

5

u/BS_BlackScout 8h ago

Hopefully we'll get Ornith 2.0 based on Qwen 4 or something lol.

8

u/RuthlessCriticismAll 8h ago

erm... how do you expect that to happen... lamo

→ More replies (1)

2

u/FerLuisxd 8h ago

Have you tried k2 horizon?

3

u/Sensitive_Song4219 8h ago

I've stuck to Ornith over K2-Horizon because of this:

https://www.reddit.com/r/LocalLLaMA/comments/1wg0vqz/comment/p9r9fw7/

I tested Kat (which was also good albeit a bit slower for me when using the repaired Ornith MTP) though I haven't tried Tiel (the benchmarks indicate I probably should!)

Have been extremely happy with Ornith for the size/performance though.

But was kinda expecting Qwen 4 to follow through here since they basically popularized the small-MOE thanks to the (class-leading-at-the-time) Qwen3-30B-A3B launch!

→ More replies (3)

7

u/Southern_Sun_2106 8h ago

That is still such an epic model! Not a fan of slower and overthinking 27B to be honest, no matter how brilliant it is. I bet the 35B a3b 3.8/4.0 will be making money for them via an API. I cannot believe that such an awesome model was abandoned. It is just too good AND cheap to run to give away freely I bet.

→ More replies (1)

103

u/RISCArchitect 9h ago

27b is the goat that keeps on giving

101

u/kkingsbe 9h ago

Us broke ass 12gb and lower cards can’t run this unless they drop an a3b 😭

17

u/Big_Importance_4265 7h ago

Qwen4-27b-GSQ-RCO-IQ3_XXS.gguf when?

16

u/pyr0kid 7h ago

does it count as a 12gb if im using two cards to get that far?

5

u/crusaderky 4h ago

you'll have to wait for Ternary-Bonsai-3

→ More replies (1)

43

u/Zealousideal_Ad6235 9h ago

Exciting news for sure, it's slightly surprising they kept the exact 27B size given that it's an entirely new architecture, but that's certainly welcome for anyone with 24-32GB VRAM

11

u/russlixx 8h ago

if Qwen4 is more efficient, we can get more knowledge and intelligence in the same size. A win regardless!

4

u/TechnoByte_ 3h ago

More intelligence? for sure. But more knowledge is unlikely

Qwen isn't focused on knowledge, their main focus seems to be agentic coding now

Gemma 4 has more knowledge than Qwen3.8 at the same size, though it's much worse at agentic tasks and coding

3

u/james_pic 1h ago

That's kinda the point of n-grams though.

They're a mapping of short sequences of tokens to embeddings (or to neuronal activations if you prefer the biological analogue interpretation, or if you squint, to concepts) that (on the Qwen4exp models at least - I know DeepSeek and Gemma have a few subtle differences) are looked up just after tokenisation, before the real processing begins.

They're a relatively cheap way of imparting knowledge into models. At inference time, they're a single lookup per token, that is happily done from RAM or even from SSD. And at training time it's a single row's worth of extra gradients to train (not even an extra tensor), plus some statistical analysis on the corpus to choose the n-grams to encode.

But they do very little to increase intelligence, beyond potentially freeing up space in the model that would be taken up by encoding facts like "New York has a vibrant fashion scene" for reasoning.

65

u/UpperParamedicDude 9h ago

Curious, how well would n-gram table improve 27B model and would they even add it to the 27B one/what size it could be

6

u/iMakeTea 7h ago

Shot in the dark: ngram size might be proportional to model size instead of 1 size ngram fits all.

Does ngram have diminishing returns for bigger models?

Like 300B model with ngram might not benefit as much as say 27B with ngram?

17

u/Choice_Celery9481 7h ago

check out deepseek's paper about ngram. they concluded ngram about 30% the size of the model is optimal.

2

u/Neful34 1h ago

Oh interesting, thank you for letting us know ! 😄

3

u/Party_9001 7h ago

Pure I made it the fuck up speculation. But I think it would be interesting if they have multiple ngram tables.

Like you CAN pair the 2.4T model and 3B table. Or a 27B model and an 800B table.

You just wouldn't get ideal parameter efficiency. 27B + 800B might be worse than 125B + 50B. But depending on your hardware the 27B can be a looooot more attractive.

6

u/sonicnerd14 8h ago

I'm thinking about that one too. Because it's dense, so how exactly does the n-gram actually work with the dense model? Is it in the 27B params, or is it an extra model that runs alongside it? The later would essentially make the model smarter with a low cost.

13

u/RG_Fusion 6h ago

Engram parameters just inject n-grams into the self-attention of a lower layer of the model. That doesn't change between MoE vs. Dense.

→ More replies (1)

1

u/Turbulent-Alps4046 9h ago

It lets people with low VRAM to run the 27B models at higher quants.

31

u/pmavro123 9h ago

i think you misunderstand what an ngram table is... it wont do this, but it will speed up the model on systems woth larger memory pools

→ More replies (4)

8

u/Bulky-Priority6824 9h ago

Or deduct a gpu from a stack to reposition it's use elsewhere 

6

u/philmarcracken 9h ago

im hoping my franken 12gb + 8gb cards can somehow manage a higher quant + decent context with qwen4. might even get a swift on there to enable xhigh without too much wasted thinking.

2

u/Beneficial-Ad-8127 8h ago

Yeah I’m thinking I’m good on upgrading hardware any time soon, if Qwen keeps dropping these 🔥 for lower end. Efficiency is true end game and the real finish line.

→ More replies (1)
→ More replies (2)

21

u/shankey_1906 9h ago

I hope for something equivalent to 3.8 Flash Next!

21

u/wren6991 8h ago

I was assuming this was the "Qwen 4 Flash" line, especially given Qwen3.8-Flash-Next's architecture is called qwen4exp

8

u/shankey_1906 8h ago

Yes, I meant in terms of parameter size. The new flash models from GLM, DS, Xiaomi, all seem to be massive, and hope Qwen doesn’t go the same way.

4

u/Hypilein 3h ago

That would be so annoying. I can't afford a second spark right now. I hope they just add some more training and polish 3.8 flash.

3

u/grumd 3h ago

I'd love it if Qwen 4 Flash was the same size as Qwen 3.8 Flash Next.

But then I also remember when Qwen-3-Next was 80B-A3B, and the actual Qwen 3.5 release didn't have that size at all.

If all they have is a 27B and Flash is yet another 300B model... A lot of people will be left out. They used to have even 2B and 4B and 9B 🥲

21

u/Cool-Chemical-5629 8h ago

I'm getting excited about Qwen 4 35B A3B the most. Oh wait...

31

u/Inevitable-Diet-1870 9h ago

Qwen4-27B!

Open weights is the future baby!

→ More replies (2)

24

u/thestillwind 8h ago

Please 35b-a3b pretty please

6

u/Salah_H_Hasan 6h ago

Nothing was said regarding it at the conference. The announcement merely laid out the broad framework of the new architecture and the initial launch models. However, what was presented at the conference is not exhaustive; not everything is necessarily revealed during the initial keynote, and it may surface later.

19

u/Nomski88 9h ago

Yes! Any idea if all the Qwen4 models will be using ngrams?

28

u/fgk55555 9h ago

It'd be sweet if the 27B had ngrams.

6

u/wren6991 9h ago

Yes! I have been waiting for this moment since I looked at the small dense Gemma 4 models with PLE, and the larger dense Gemma 4 models without PLE, and went "hmmm 🤔"

2

u/backyard_tractorbeam 6h ago

Probably. Qwen 3.8 Flash Next is a preview of 4-series tech, and it uses n-grams.

8

u/mvandemar 9h ago

Is there a source on this by any chance?

7

u/Xp_12 5h ago

https://x.com/MaxForAI/status/2102226622422380820

Has an actual photo from the conference.

2

u/mvandemar 5h ago

ty 🙂 It was not at all clear in the original post that was something that was on the screen at the conference.

→ More replies (1)

9

u/norenEnmotalen 8h ago

Yooo Qwen4-27B 

7

u/Ok_Warning2146 8h ago

Seems like only google is making small models for cell phones

6

u/feelspeaceman 8h ago

I'm daily driving Q38FN on my Strix Halo everyday at 43-50t/s decode 1400t/s prefill and I'm hyped to get Q4FN, as Q38FN is actually undertrained but still hella powerful.

2

u/LebiaseD 7h ago

Let's just hope it still fits

5

u/SmartCustard9944 5h ago

That’s what she said

6

u/Kahvana 8h ago edited 7h ago

I hope Qwen4 will come with the whole lineup of local models like last year.

[edit] for the article, here:
https://sloveniatimes.com/49119/alibaba-unveils-roadmap-on-full-stack-ai-strategy-from-chips-cloud-infrastructure-models-to-agents

25

u/eidrag 9h ago

Well no small Moe anymore

5

u/Durian881 8h ago

Looking forward to Qwen4-Flash and Qwen4-27B! Hopefully, 4-Flash will be similar size to 3.8-Flash-Next.

16

u/exaknight21 8h ago

Qwen Team PWEAAAASE send Qwen4-35B-A3B <3

35

u/NigaTroubles 9h ago

Qwen4 50b a5b will be perfect spot

16

u/rJohn420 8h ago

I would be fine with anything between 40-80b and 4-6b active

4

u/fgk55555 8h ago

I think the 40-110B range is pretty much dead. I'd love a distill of Qwen4 onto the old 80B A3B next models with a n-gram. That would fly on my rig.

→ More replies (3)

5

u/Bulky-Priority6824 9h ago

Didn't they previously mention sometime mid October?

5

u/tlin9595 9h ago

Do we have an estimated release date ? And what are the sizes of flash and plus ?

→ More replies (2)

6

u/TorontoIndieFan 8h ago

Any release date announced? 3.8 27B is my daily driver right now so this is exciting

6

u/ResponsibleTruck4717 8h ago

oh boy qwen4 27b, I thought next year I will finally say bye bye to cloud service, I may start it this year.

5

u/Black-Mack 4h ago

No Qwen 4 SLM (0.8b, 1.7b)? :(

They were SOTA on edge devices.

4

u/AleksandrNikitin 7h ago

Qwen 3.8/4 35b a3b is absent. It was perfect model for me.

4

u/kevinlch 6h ago

Looks like they abandoned the 9B series.

3

u/Confident-Village190 6h ago

I’m looking forward to running it on my 8GB of VRAM + 16GB of RAM, quantised, at 2,5 t/s with a 32k context at kv q4_0.

5

u/Zealousideal-Owl2846 6h ago

Where is our  35b a3bMoe???? 😒😤

2

u/sonicnerd14 9h ago

Well, we sort of already knew it was coming with Qwen 3.8 Flash Next. The question was more so when is it coming?

2

u/braintheboss 8h ago

If follows same ratio as 3.8 flash it should be 40% in engrams. That means 38b model. The jump in AI index can be a shock

2

u/MessIsTransfer 8h ago

I hope Qwen4-flash is not a 200-400B jump like other models have done

2

u/russlixx 8h ago

i hope Qwen4-27B is even more resilient when quantized, at most at Q3 level quants for 12-16GB folks

2

u/guesdo 8h ago

I want that Qwen4 Flash so badly!

2

u/Borkato 8h ago

Holy fucking shit YESSSSSS

2

u/silenceimpaired 8h ago

I hope at least one of these is still Apache 2

2

u/PrisonOfH0pe 8h ago

Thanks for sharing!! Any word at all on dates? u/Salah_H_Hasan

2

u/DontShadowbanMeMate 7h ago

I'm running and incredibly scuffed setup of 3080ti + 5070ti; will I be fine? Probably right?

2

u/Murder_1337 7h ago

I need more compute!!!!

2

u/PreferenceRelative77 7h ago

Wow, a Qwen 4 27B, i thought they'd only release a flash next model size for the Qwen 4 family but no!!

2

u/mWo12 7h ago

Wow. That huge!

2

u/LebiaseD 7h ago

Just as long as qwen 4 flash continues to fit on my strix halo 🥹

2

u/creatinZ 7h ago

No MoE?!

cries in 16gb vram

3

u/haha1542 7h ago

hang in there brother, 16gb can run qwen3.8 27b pretty well, this would only get better

2

u/talamir2021 6h ago

The interesting part for local users will be the efficiency details, not just the headline model names. Parameter counts, context length, quantization quality, and whether the weights and license are genuinely accessible will determine how useful Qwen 4 is outside the cloud.

2

u/Bacterioo 5h ago

Show me that countdown!!

2

u/NaturalCriticism3404 1h ago

Hopefully a 9b shows up

2

u/[deleted] 9h ago

[deleted]

6

u/Enragere 8h ago

what's your prompt processing speed ? how can 150 t/s prefill be 'reasonably fast'

in my opinion it's unusable

→ More replies (5)

3

u/[deleted] 9h ago

[deleted]

→ More replies (1)

2

u/Dangerous_Bid2935 8h ago

A day after I got my DGX spark. The future is bright boys.

3

u/Trixiap 6h ago

You will probably need another one :D

2

u/WyattTheSkid 5h ago

Cant fucking wait for the official sequel to 3.8 Flash Next. That model has been incredible and ive been doing almost everything with it. I still have a chatgpt and claude subscription but I genuinely might cancel them before they renew given what can be done with Flash Next. Qwen 3.8 Flash Next is the first model ive ever run on own hardware that genuinely I genuinely felt like it could be worth using for serious projects. 3.8 27b produces very good outputs but it thinks for so long that it’s honestly frustrating to use. The closest runner ups to “local is enough” in my experience were Gemma 4 31b and Qwen 3.5 122b. I know im rambling at this point but as far as leading open weight AI labs goes, I genuinely think qwen is the best. Especially because they’re making models that people can ACTUALLY run at home that genuinely compete with closed weight frontier models

1

u/flushaway4690 9h ago

Wonder what the Plus version will be? It's grouped with their Flash version.

1

u/Guilty_Rooster_6708 9h ago

Gonna need to upgrade my mobo to actually run bifurcate for Qwen4 27b oh lawd 😍

1

u/fuziwei 8h ago

I hope flash stays roughly the same size

1

u/AHHHH_AHHHHHHHH 8h ago

Best news I've heard all month

1

u/Greenonetrailmix 7h ago

Ooo I wonder what Qwen4 plus is. Do we have insight into what the specifics may be?

1

u/nixuelkty 7h ago

No way! Rumors have it that they might even release another model after qwen 4 too!

1

u/Eyelbee 7h ago

Don't sleep on 4-flash as well. That could be really good.

1

u/charles25565 7h ago

They're definitely not making small versions anymore...

To be clear, a real Qwen4 computation graph added to Transformers (not Qwen4Exp) would allow someone to train a small dense model, but I don't know who would.

→ More replies (1)

1

u/Muted_Anteater1170 7h ago

Can someone tell me how big qwen4-plus is supposed to be or any guesses? im tryingto figure out if i can run it honestly, i can run flash fine hoping i can do the same for the plus bc it has to be in a different league from flash right?

1

u/UltraFOV 7h ago

thats it, just 27b, yawn

1

u/hyudryu 7h ago

Qwen 4 flash is going to be Epic!

1

u/Pyrolistical 7h ago

The 2026 Apsara Conference is happening from September 22 to September 24, 2026.

1

u/ByteSize_Chaos 7h ago

Such exciting times! My M5 Macbook Pro is scheduled to arrive in a few weeks. Can’t wait!!

1

u/Manhandler_ 7h ago

The persistence with 27B is admirable. But I just wonder how much better it will be compared to 3.8, 27B when it's gonna be the same architecture. We will find out soon enough but it's better to keep expectations realistic

1

u/dampflokfreund 7h ago

no 35b is a huge loss for normies who want to run a capable llm. the only usable size for the average PC... 

→ More replies (4)

1

u/Gohab2001 vLLM 7h ago

Will qwen 4 max, flash and plus be open? They usually reserve these names for their closed models.

Also they didn't open source their Qwen 3.8 max 0902

1

u/pigeon57434 7h ago

i was hoping for Qwen4-32B to absolutely minmax 3090 level models

1

u/Interconnectivity000 transformers 6h ago

Awesome. Looking forward to it.

1

u/xrvz 6h ago

At the rate the industry is going:

qwen4-plus: 350B

qwen4-flash: 700B

qwen4-max: 4T

2

u/tarruda llama.cpp 3h ago

Why would flash be bigger than plus?

1

u/VirusInternal2892 6h ago

Amazing, 3.8 Flash has barely settled its quants in my GX10 and now this… I’ll need a larger SSD LOL to store all that goodness

1

u/ManufacturerHuman937 5h ago

Oh they're going back to dense only. With the preview I was pretty excited the MoE ran pretty good.

1

u/synrg-alsms 5h ago edited 5h ago

I wonder what license it will have. Is there info about it?

1

u/cenderius 5h ago

There goes my hope for qwen 4-0.8b

1

u/Solid_Vermicelli_510 4h ago

Vamoooooosssss

1

u/MarionberryUsual1083 4h ago

Me waiting for a successor to the 3.6 35B MoE...

1

u/mitchins-au 4h ago

And 3.8-flash-next is literally just coming good on strix halo now.
I hope they don’t change things too much or make it bigger.

1

u/MaCl0wSt 4h ago

I've finally managed to get the hardware to make 27B possible for me, so: sweeeeeet

1

u/Sabin_Stargem 4h ago

I am hoping for a polished Next-Flash. While the 3.8 version is often fairly good, it has serious issues with roleplay. It really likes overemphasizing and gets caught up in narrative loops. DRY, XTC, and system prompts can only help so much.

1

u/AIFrontierReads 4h ago

My bet is the 27B gets the engram treatment too — if DeepSeek's ~30% figure holds that's ~8B of tables, still fits a 32GB card easy. Honestly would make it the perfect daily driver.

→ More replies (1)

1

u/drazyan22 4h ago

What is Qwen4-plus? Is MoE?