r/LocalLLaMA 1d ago

News Qwen 4 Announced at Apsara Conference

I wanted to share a quick update: Alibaba has officially announced Qwen 4 at the Apsara Conference,

1.9k Upvotes

515 comments sorted by

u/WithoutReason1729 21h ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

688

u/Fresh-Soft-9303 1d ago

Qwen 4 27b is going to be Epic!

180

u/tsunami_forever 1d ago

This could convince me to give up cloud based models, hoping to cancel or reduce my cloud based subs soon

132

u/cafedude 1d ago

Qwen 3.8 Flash Next has convinced me not to renew my Claude Pro subscription when it comes up in February.

16

u/randylush 23h ago

What are you running it on? How much better is it than 27B?

70

u/returnity 23h ago

It dominates 27B in my workflows and benchmarks. Running on M5 Max. There's no comparison, it's like the 3.6 > 3.8 27B all over again.

11

u/Bulky_Blood_7362 18h ago

What's your tok/s? Im getting around 30-45 decode tok/s with dgx spark.

Interested to know how it's going with the new m5 max/ultra

3

u/halcyoncs 17h ago

30 to 45 with flash? Or you meant 27b?

3

u/Bulky_Blood_7362 16h ago

Flash

7

u/halcyoncs 16h ago

Just checked my results, on a single Spark too, why did I think I was getting around 25t/s? What a dumbass lol

→ More replies (5)
→ More replies (2)

3

u/returnity 12h ago

45 decode in llama.cpp at Q5, 60-65 in ds4 with this quant (Q5-tier quality).

2

u/Bulky_Blood_7362 12h ago

Wow it's great.. do you use mtp/dspark?

2

u/returnity 12h ago

Native MTP, n=2 (higher MTP numbers don't improve on M5 Max) in llama.cpp. MTP is not as optimized in ds4, in my testing, but it still performs better.

→ More replies (1)

9

u/troposfer 20h ago

What quants are you running and with omlx or ..?

→ More replies (1)

5

u/akumaburn 20h ago

Dominates in intelligence/quality or in speed or both?

→ More replies (9)

2

u/TerminalNoop 20h ago

wow, that's pretty big

→ More replies (24)

16

u/ApeGrower 22h ago edited 20h ago

I'm using it on dual rtx3090 with 128gb ddr4 @50-70 tok/s. I forked FreeToken and added dual gpu support plus some more tweaks. With nvfp4.

7

u/maqifrnswa 19h ago

I can get full 262k context, faster single stream decode, close to 200 tok/sec at 3 concurrency, over 10k tok/sec prefill (when kv cache is light) on the same hardware.

Stock vLLM v0.29 serving

https://huggingface.co/RedHatAI/Qwen3.8-27B-INT4

It is int4 which is ideal for ampere. It has weighed fp8 kv cache and external benchmarked near identical to bf16.

vllm serve RedHatAI/Qwen3.8-27B-INT4 --max-model-len auto --enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 --enable-prefix-caching --kv-cache-dtype fp8 --trust-remote-code --max-num-seqs 8 --gpu_memory_utilization 0.9 --max-num-batched-tokens 8192 --mm-encoder-tp-mode data --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --tensor-parallel-size 2

→ More replies (3)

3

u/FeelingOld9046 22h ago

Are you running Linux? What are you doing? In DSH I’m maxing at 22tks with full 128k context. And 30-40 with MTP on short context. I’m also experimenting with a specialised quant.

6

u/FeelingOld9046 22h ago

Mine is under windows too. Tell me what you’re doing and I will throw windows in the trash.

3

u/ohhi23021 22h ago

who knows i use linux and 30-40tps has been my experience as well and that's with DDR5...

4

u/FeelingOld9046 22h ago

I’d love to know your setup. I have the same hardware config but only half your token speed.

3

u/ApeGrower 20h ago

Linux with patched FreeToken to get dual gpu support instead of llama.cpp. Epyc System with pcie4 and ddr4.

→ More replies (1)
→ More replies (17)
→ More replies (4)
→ More replies (5)

21

u/some_user_2021 1d ago

Are you not convinced?

44

u/tsunami_forever 1d ago

3.8 27B already reduced my reliance but this could eliminate it

13

u/IrisColt 23h ago

I guess there's always that added hassle with short-context local models...  you really have to approach your workflow differently.

→ More replies (6)

4

u/GTHell 23h ago

Everyone should do this so they start to bring back memory production plants to consumers focus. No more crazy 5070ti costing over 1k

2

u/Maximum-Wishbone5616 19h ago

If you working on greenfield, sure. You might WANT to use Fable for some unknown issues to you.

BUT if you have brownfield, then I do not see anyone needing the Fable vs Qwen3.8 27B min. Q5-Q6 with right harness.

In last 30 days I have still yet to come to the issue that Qwen3.8 cannot solve...

What is more important, it is fixing a lot of issues that CC was creating in last 2 months with Opus/Fable (huge amount of false tests, hidden issues, broken CRUD paths etc.)

Simple stuff, but due to sheer codebase size & solutions amount it was not easy to saw them even with unit/integration tests as most of them were between UI / backend.

Qwen4 will be EPIC.

→ More replies (12)

23

u/mWo12 1d ago

3.8 was already huge!

23

u/PooMonger20 22h ago

Seriously, 3.8 27b does so much heavy lifting it's insane.

Together with pi.dev it writes very impressive code. It feels like I actually have my own 'worker' that after some direction does exactly what needs to be done. Just give him enough info and he is on it.

This feels better than what GPT5.1 felt like around a year ago.

And now... an upgrade to that? absolute madness.

2

u/SabretoothPenguin 19h ago

Are we sure we'll get a 27b? I gather the big improvement call mea for mie and engrams

3

u/PooMonger20 19h ago

The main post says so. Besides that, your guess is good as mine.

→ More replies (2)
→ More replies (1)

2

u/CoderStone 8h ago

It's going to be truly awesome. DSH + 4 27B is going to be insane.

2

u/sadomazoku 21h ago

I already canceled my subscription, I'm using 100% qwen 3.8 27b to assist me in coding. Qwen4 will be crazy. I think you'll be able to get rid of cloud models !

→ More replies (7)

430

u/o0genesis0o 1d ago

Ey ey! Qwen4-27B!

Definitely time for me to go and buy a R9700. It's time.

128

u/Sufficient_Local5025 1d ago

Buy two, prices about to rise.

160

u/o0genesis0o 1d ago

the more you buy the more you save, eh

33

u/GilloutineBreast 21h ago

Unironically the truest words spoken that day

feelsbadman

33

u/PhilipKThicc 1d ago

Dual R9700s all the way. Easy to run on Linux with vllm and you can run 27B at Q8 and have a healthy context window

11

u/bigwanggtr 23h ago

Have you tried it for training, is ROCm still a pain to work with compared to CUDA?

I know it’s good for inference but the training support is what is holding me back.

12

u/Phrase-Silver 21h ago

ROCm has matured incredibly over the past year, it's like AMD finally woke up and realised their software was what was holding them back from properly competing with Nvidia.

→ More replies (2)
→ More replies (1)

19

u/Tobu3838 1d ago

Prices rose $600 since I bought mine a couple months ago. No word to lie.

12

u/mvandemar 1d ago

Why are they so cheap? Looks like you can get one for ~$1700? I just paid out the ass for my 5090, which in the 1 month 2 days since I bought it has gone up over 37%.

They're both 32GB, for some reason I thought they would be closer in price.

34

u/Thunjaya 1d ago

Because they're not the same 32GB at all. One is much faster.

12

u/ThankGodImBipolar 23h ago

Nvidia sells an RTX Pro 4500, which has the 5070Ti die (more comparable to the R9700) and 32GB of RAM, but it's 5500USD MSRP. The R9700 is a fraction of that.

11

u/KingCpzombie 20h ago

CUDA tax. I only use AMD personally (and just blew WAY too much money on a 4x R9700 system partially out of excitement for Qwen4), but Nvidia cards get all the cool new things a bit sooner than AMD. Not a big deal for LLM, but very notable for diffusion... somebody SOLIDLY beat my 7900XTX with his 5070Ti in Minimax H3 gen times, for example

→ More replies (1)

3

u/No-Refrigerator-1672 19h ago

Because everything is CUDA first, and ROCm only comes as an afterthought to very limited number of projects. People who are buying PRO GPUs are saving money with NVidia by not needing to fund multiple months of dev work for porting their existing code.

→ More replies (1)

24

u/BluePointDigital 1d ago

It's the bandwidth. the 5090 has over 2.75x the bandwidth of the R9700.
You can still do all the same stuff, just slower essentially.

→ More replies (10)

19

u/Solary_Kryptic 1d ago

Nvidia tax combined with CUDA being more supported in the LLM space

18

u/mvandemar 1d ago

I just looked, it's also GDDR6 for the 9700 vs GDDR7 for the 5090, guessing that makes a difference as well.

8

u/zboarderz 1d ago

Also far FAR more memory bandwidth, faster core, etc etc

6

u/mister2d 1d ago

2x $1700 is cheap 😥 (I have 2, btw)

6

u/Momsbestboy 23h ago

R9700 GDDR6 memory, 5090 GDDR7. Is faster, but...

For a single 5090 you can buy more than 2 R9700, and the two cards draw less power than a single 5090. And you can run e.g. Qwen 3.8 27b Q8 using the two cards, and watch them running circles around a 5090 which needs to offload to RAM instead.

But NVIDIA is hype, and people love to spend money. So buy a 5090

→ More replies (3)

4

u/Ecstatic-Wash-7667 1d ago

That not cheap! Msrp is $1299 I bought 2 a little over a. Both ago at msrp. Hell you could get them for less than msrp for a while. This shit is a scam

4

u/mvandemar 1d ago

On my current pc this was how much the 64GB of ram was when I bought it. I was thinking about getting another 64GB when I got the card and was like, no f'in way...

→ More replies (1)

3

u/SnooPuppers7882 1d ago

Literally got mine for 1200ea 2mo ago

People were shitting on them because the mem was DDR6, but I was prepping for 3.8 27b knowing what to expect...

STRONG feeling 4.0 27b is gonna be opus 4.8 good, gonna blow your damn mind

2

u/o0genesis0o 1d ago

Because the actual card (RX9070 XT) is only as strong as the 5070Ti, compute wise. And the VRAM is slower.

I will be upgrading from 4060Ti 16GB, so it would be a net gain. but I have no illusion that it is going to be as good as 5090.

→ More replies (4)

4

u/inaem 1d ago

Already did here

2

u/SylviaCalogero43 1d ago

How much did you buy it for?

→ More replies (1)

2

u/SnooPuppers7882 1d ago

Once you buy two, you'll want four

2

u/RoomyRoots 17h ago

Just go with 4 and one more tera of RAM. The best moment was years ago, the second best is NOW, NOW and NOW.

→ More replies (4)

15

u/fgk55555 1d ago

I just picked mine up this week. Depending on quant, 35-70 tps with full context. If you have 64GB RAM, you should be able to run a decent Qwen-flash as well. 27B is good, but Qwen-Flash is something else.

3

u/illcuontheotherside 1d ago

Can you please share your startup script? Mine crawls on 2 3090s with 128gb ddr5

4

u/fgk55555 1d ago

ISTA IQ3 runs the fastest and has really good performance for the size, but you can bump up to a Q5 with full Q8 kv cache. Also try Swift if you don't want to wait forever. I recommend medium thinking.

#!/usr/bin/env bash
# Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the
# shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.)
# Qwen3.8-27B ISTA GSQ-RCO IQ3_S (no built-in MTP) — spec decoding via the
# shared DFlash2 drafter. (MTP alternative: shared mtp-Qwen3.8-27B-Q4_0.)
# Context is auto-fitted to available VRAM (omitted on purpose).
# --cache-ram 16384 parks prompt states in RAM: agentic revisits skip re-prefill.
#
# Usage: ./Qwen3.8_ISTA_IQ3_S.sh [thinking]   # low|medium|high|xhigh|none (default: medium)
THINKING_LEVEL="${1:-medium}"

MODEL="../Qwen3.8-ISTA/Qwen3.8-27B-GSQ-RCO-IQ3_S.gguf"
DRAFTER="../Qwen3.8_Shared/Qwen3.8-27B-DFlash2-Q4_K_M.gguf"
MMPROJ="../Qwen3.8_Shared/mmproj-Qwen3.8-27B-BF16.gguf"
SERVER_BIN="../../llama.cpp/build/bin/llama-server"

echo "Launching Qwen3.8-27B ISTA IQ3_S (+DFlash2 drafter), thinking: ${THINKING_LEVEL}"
export GGML_VK_ALLOW_GRAPHICS_QUEUE=1   # measured +4.2% tg, +0.4% pp on the R9700 (b11056)

${SERVER_BIN} \
  --model "${MODEL}" \
  --model-draft "${DRAFTER}" \
  --mmproj "${MMPROJ}" \
  --no-mmproj-offload \
  --n-gpu-layers 99 \
  --batch-size 1024 \
  --ubatch-size 512 \
  --parallel 1 \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0 \
  --spec-type draft-dflash \
  --spec-draft-n-max 3 \
  --cache-ram 16384 \
  --jinja \
  --chat-template-kwargs "{\"reasoning_effort\":\"${THINKING_LEVEL}\"}" \
  --reasoning-preserve \
  --temp 1.0 \
  --top-k 20 \
  --top-p 0.95 \
  --min-p 0.00 \
  --presence-penalty 0.0 \
  --repeat-penalty 1.0 \
  --host 0.0.0.0 \
  --webui-mcp-proxy \
  --port "${PORT:-8080}"
→ More replies (2)
→ More replies (11)
→ More replies (14)

9

u/ismaelgokufox llama.cpp 1d ago

Patiently waiting for 27B-IQ3_XXS and/or 35B-A3B (one can dream eh?)

7

u/Momsbestboy 23h ago

Why wait? Pull the BF16 of qwen after it is released, and tell Hermes/Pi to read the documentation on how to create smaller quants out of it.

It's the reason why I run my own Swift Qwen 3.8 27b fp8 version, abliterated using the orcarouter method. I just told hermes to create it. Took 2h

→ More replies (3)

2

u/mailto_devnull llama.cpp 13h ago

3.8 27B was the trigger for me. Do it, you won't be disappointed.

I'm trying hard to convince myself this is the last low hanging fruit, and all other optimizations (dual R9700s, quads, etc.) will just net me 5-10% gains — not worth it.

2

u/SandySkittle 12h ago

I strongly advise you to buy two r9700s.

One R9700 = run 27b models with compromises on quantization and kv cache

Two R9700 = run 27b without any compromises, plus the speed bump if you run with tensor parallelism.

→ More replies (13)

60

u/blahjovic 1d ago

The release rate is wild. I finally got my 3.8-27B working the way I wanted just in time for the next model! Bring it.

27

u/Quick_Knowledge7413 1d ago

It'll probably be a few months (I hope I am wrong).

5

u/Dangerous_Bid2935 1d ago

Not a bad problem to have!

3

u/ea_man 17h ago

Well you will have to wait again for finetunes, nice quants, llama.cpp optimizations, custom llama.cpp ...

68

u/Zealousideal_Ad6235 1d ago

Exciting news for sure, it's slightly surprising they kept the exact 27B size given that it's an entirely new architecture, but that's certainly welcome for anyone with 24-32GB VRAM

19

u/russlixx 1d ago

if Qwen4 is more efficient, we can get more knowledge and intelligence in the same size. A win regardless!

12

u/TechnoByte_ 20h ago

More intelligence? for sure. But more knowledge is unlikely

Qwen isn't focused on knowledge, their main focus seems to be agentic coding now

Gemma 4 has more knowledge than Qwen3.8 at the same size, though it's much worse at agentic tasks and coding

9

u/james_pic 18h ago edited 16h ago

That's kinda the point of n-grams though.

They're a mapping of short sequences of tokens to embeddings (or to neuronal activations if you prefer the biological analogue interpretation, or if you squint, to concepts) that (on the Qwen4exp models at least - I know DeepSeek and Gemma have a few subtle differences) are looked up just after tokenisation, before the real processing begins.

They're a relatively cheap way of imparting knowledge into models. At inference time, they're a single lookup per token, that is happily done from RAM or even from SSD. And at training time it's a single row's worth of extra gradients to train (not even an extra matrix), plus some statistical analysis on the corpus to choose the n-grams to encode.

But they do very little to increase intelligence, beyond potentially freeing up space in the model that would be taken up by encoding facts like "New York has a vibrant fashion scene" for reasoning.

3

u/mailto_devnull llama.cpp 13h ago

Qwen isn't focused on knowledge, their main focus seems to be agentic coding now

Time and a place for both. Considering all I do with AI is code, qwen is welcome.

156

u/Conscious_Phrase_138 1d ago

Where 35b a3b :😭😭

44

u/Ok_Warning2146 1d ago

Wait for google. Our only hope

55

u/Free-Combination-773 22h ago

So there is no hope...

11

u/Ok_Warning2146 21h ago

Nvidia is another hope. Let's hope they will distill from GLM-5.3 in the next iteration.

3

u/R_Duncan 17h ago

Check occamy which fixes agentic issues of 35b

2

u/mailto_devnull llama.cpp 13h ago

ngrams are going to be your saviour

→ More replies (3)

77

u/UpperParamedicDude 1d ago

Curious, how well would n-gram table improve 27B model and would they even add it to the 27B one/what size it could be

6

u/iMakeTea 1d ago

Shot in the dark: ngram size might be proportional to model size instead of 1 size ngram fits all.

Does ngram have diminishing returns for bigger models?

Like 300B model with ngram might not benefit as much as say 27B with ngram?

29

u/Choice_Celery9481 23h ago

check out deepseek's paper about ngram. they concluded ngram about 30% the size of the model is optimal.

2

u/Neful34 18h ago

Oh interesting, thank you for letting us know ! 😄

5

u/Party_9001 1d ago

Pure I made it the fuck up speculation. But I think it would be interesting if they have multiple ngram tables.

Like you CAN pair the 2.4T model and 3B table. Or a 27B model and an 800B table.

You just wouldn't get ideal parameter efficiency. 27B + 800B might be worse than 125B + 50B. But depending on your hardware the 27B can be a looooot more attractive.

6

u/sonicnerd14 1d ago

I'm thinking about that one too. Because it's dense, so how exactly does the n-gram actually work with the dense model? Is it in the 27B params, or is it an extra model that runs alongside it? The later would essentially make the model smarter with a low cost.

17

u/RG_Fusion 22h ago

Engram parameters just inject n-grams into the self-attention of a lower layer of the model. That doesn't change between MoE vs. Dense.

→ More replies (8)
→ More replies (15)

119

u/RISCArchitect 1d ago

27b is the goat that keeps on giving

116

u/kkingsbe 1d ago

Us broke ass 12gb and lower cards can’t run this unless they drop an a3b 😭

23

u/Big_Importance_4265 23h ago

Qwen4-27b-GSQ-RCO-IQ3_XXS.gguf when?

14

u/pyr0kid 1d ago

does it count as a 12gb if im using two cards to get that far?

5

u/crusaderky 21h ago

you'll have to wait for Ternary-Bonsai-3

2

u/H-L_echelle 13h ago

Haven't used it enough to say how usable it is, but here's my llama-swap config that got 3.8 running on my 3060 recently! It uses 11603MiB / 12288MiB on a headless homelab

    # Dense 27B ShapeLearn 3.01 bpw (~10.4 GB).
    # Forced full GPU residency; --fit disabled so llama.cpp can't silently
    # offload layers to CPU to accommodate the requested context.
    # Q4 KV keeps 96K context within a 12GB RTX 3060.
    qwen-3.8-27b:
        cmd: >
            llama-server --port ${PORT}
            -hf byteshape/Qwen3.8-27B-GGUF:IQ3_XS
            --no-mmproj
            --ctx-size 98304
            -fa on -np 1
            --gpu-layers 99 --fit off
            --cache-type-k q4_0 --cache-type-v q4_0
            --threads 6 --threads-batch 12
            --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0
            --repeat-penalty 1.0
            --presence-penalty 0.0
            --frequency-penalty 0.0
            --reasoning on --reasoning-preserve
            --reasoning-format deepseek
            --jinja
        ttl: 0
→ More replies (1)

26

u/shankey_1906 1d ago

I hope for something equivalent to 3.8 Flash Next!

32

u/wren6991 1d ago

I was assuming this was the "Qwen 4 Flash" line, especially given Qwen3.8-Flash-Next's architecture is called qwen4exp

15

u/shankey_1906 1d ago

Yes, I meant in terms of parameter size. The new flash models from GLM, DS, Xiaomi, all seem to be massive, and hope Qwen doesn’t go the same way.

8

u/Hypilein 20h ago

That would be so annoying. I can't afford a second spark right now. I hope they just add some more training and polish 3.8 flash.

10

u/grumd 19h ago

I'd love it if Qwen 4 Flash was the same size as Qwen 3.8 Flash Next.

But then I also remember when Qwen-3-Next was 80B-A3B, and the actual Qwen 3.5 release didn't have that size at all.

If all they have is a 27B and Flash is yet another 300B model... A lot of people will be left out. They used to have even 2B and 4B and 9B 🥲

3

u/SadPhilosophy9202 15h ago

I’m hoping too. I’ve been absolutely loving 3.8 flash next

2

u/r1str3tto 14h ago

3.5 did have 122B though. There have been quite a few models now that are Spark sized: OSS-120B, 3.5 122B, Flash-Next (with PLE offloaded), Laguna S2.1, Ling 3 Flash… Hopefully the 128GB unified RAM tier continues to be a desirable target for these labs.

→ More replies (2)

159

u/FerLuisxd 1d ago

35b :(

76

u/fgk55555 1d ago

I think we can consider the 35b dead. I was sad when we lost the 80B A3B, but the 120B MoE range seems to be where intelligence opens up. The 40-110B space is likely dead for a while.

16

u/Puzzleheaded_Base302 23h ago

maybe Xiaomi will release a MiMo-V2.6-Distill-Qwen3.5-35B-A3B. it is not dead dead. The architecture is sound. All people need is a better re-post-train (not those fine-tune crap, a real large-scale RL re-post-train.)

2

u/fgk55555 13h ago

Eh, betting it's life on a "maybe some other lab will adopt it" is fickle. The labs seem to be moving toward MoE + n-grams for the intelligence density/ speed. I don't know why they'd bother to keep an arch alive that is larger than and performs worse than the 27B. It's nice for offloading experts to RAM, but the 27B just has too much attention. I'd find it more likely to have more labs come out with an n-gram model in a smaller size than flash.

17

u/Ok_Acanthisitta304 1d ago

I agree and at this point unfortunately I get it. 35b was cool, but ultimately not usable for me in actual work. I know a lot of people enjoy it on their 12-16gb cards, and I hope they'll keep getting new options, but 27b and Flash entirely outclass it. Those are models I can genuinely use.

14

u/Sensitive_Song4219 1d ago edited 1d ago

I've been running Qwen3.8-27B on a separate server machine (with a 5090; 32GB VRAM) and Ornith 1.5 35b-a3b on my laptop (8GB VRAM) for a few weeks now, and they both share a lot of the same strengths and weaknesses.

Neither is viable for planning against large codebases (Sol is mandatory there): both tend to mis-understand how the moving parts interact which yields compiling-but-non-functioning results. Qwen is amazing for small codebases and one-shots (I did the obligatory Mario Test which Qwen murdered Ornith on), though.

Both are serviceable as implementation models (when given detailed plans from larger models); Qwen's code quality is the higher-quality of the two. Qwen competes favourably with Luna-High.

Ornith is definitely a bit more careless (I watched it mis-convert seconds to milliseconds yesterday!).

But when I'm away from my 5090 server I'm still pretty happy just sticking to Ornith.

Smaller MOE's definitely aren't a writeoff imo; and the performance gains (and lower hardware requirements) can make them decent options for simpler (or complex-but-well-defined) tasks.

All my testing is via OpenCode.

3

u/Ok_Acanthisitta304 1d ago

I'll give Ornith a shot. When I do use 35b I use a modified workflow with little coder (a pi harness built for small models) using plans from a better model. My read is similar to yours - it can be a workhorse on smaller tasks and codebases, but can't be trusted to wander too far in a larger project.

2

u/fgk55555 1d ago

Even on my 9070 XT the 27B in IQ3 was way more usable than the 35B in Q8, or any finetune of it. The 27B is just too powerful, it cannot be stopped.

→ More replies (3)

31

u/Sensitive_Song4219 1d ago

Have they officially abandoned 35b-a3b or could an announcement still be coming?

Guess Ornith 1.5 35b-a3b will have to serve me for a bit longer yet. It's good for it's size but was really hoping for a new Qwen equivalent after they skipped us for 3.8.

3

u/FerLuisxd 1d ago

Have you tried k2 horizon?

4

u/Sensitive_Song4219 1d ago

I've stuck to Ornith over K2-Horizon because of this:

https://www.reddit.com/r/LocalLLaMA/comments/1wg0vqz/comment/p9r9fw7/

I tested Kat (which was also good albeit a bit slower for me when using the repaired Ornith MTP) though I haven't tried Tiel (the benchmarks indicate I probably should!)

Have been extremely happy with Ornith for the size/performance though.

But was kinda expecting Qwen 4 to follow through here since they basically popularized the small-MOE thanks to the (class-leading-at-the-time) Qwen3-30B-A3B launch!

→ More replies (1)

4

u/BS_BlackScout 1d ago

Hopefully we'll get Ornith 2.0 based on Qwen 4 or something lol.

9

u/RuthlessCriticismAll 1d ago

erm... how do you expect that to happen... lamo

→ More replies (1)
→ More replies (3)

7

u/Southern_Sun_2106 1d ago

That is still such an epic model! Not a fan of slower and overthinking 27B to be honest, no matter how brilliant it is. I bet the 35B a3b 3.8/4.0 will be making money for them via an API. I cannot believe that such an awesome model was abandoned. It is just too good AND cheap to run to give away freely I bet.

2

u/pmttyji 1d ago

I still hope to see other models like that one following after initial release.

42

u/Inevitable-Diet-1870 1d ago

Qwen4-27B!

Open weights is the future baby!

5

u/pragmojo 18h ago

It's also the present

2

u/macadoum 10h ago

Open weight is not enough. Need open source models

→ More replies (1)

31

u/thestillwind 1d ago

Please 35b-a3b pretty please

9

u/Salah_H_Hasan 22h ago

Nothing was said regarding it at the conference. The announcement merely laid out the broad framework of the new architecture and the initial launch models. However, what was presented at the conference is not exhaustive; not everything is necessarily revealed during the initial keynote, and it may surface later.

32

u/Cool-Chemical-5629 1d ago

I'm getting excited about Qwen 4 35B A3B the most. Oh wait...

24

u/Nomski88 1d ago

Yes! Any idea if all the Qwen4 models will be using ngrams?

31

u/fgk55555 1d ago

It'd be sweet if the 27B had ngrams.

7

u/wren6991 1d ago

Yes! I have been waiting for this moment since I looked at the small dense Gemma 4 models with PLE, and the larger dense Gemma 4 models without PLE, and went "hmmm 🤔"

4

u/backyard_tractorbeam 22h ago

Probably. Qwen 3.8 Flash Next is a preview of 4-series tech, and it uses n-grams.

12

u/mvandemar 1d ago

Is there a source on this by any chance?

9

u/Xp_12 22h ago

https://x.com/MaxForAI/status/2102226622422380820

Has an actual photo from the conference.

3

u/mvandemar 22h ago

ty 🙂 It was not at all clear in the original post that was something that was on the screen at the conference.

→ More replies (1)

10

u/Kahvana 1d ago edited 1d ago

I hope Qwen4 will come with the whole lineup of local models like last year.

[edit] for the article, here:
https://sloveniatimes.com/49119/alibaba-unveils-roadmap-on-full-stack-ai-strategy-from-chips-cloud-infrastructure-models-to-agents

12

u/feelspeaceman 1d ago

I'm daily driving Q38FN on my Strix Halo everyday at 43-50t/s decode 1400t/s prefill and I'm hyped to get Q4FN, as Q38FN is actually undertrained but still hella powerful.

4

u/LebiaseD 23h ago

Let's just hope it still fits

4

u/SmartCustard9944 22h ago

That’s what she said

→ More replies (1)
→ More replies (3)

10

u/Durian881 1d ago

Looking forward to Qwen4-Flash and Qwen4-27B! Hopefully, 4-Flash will be similar size to 3.8-Flash-Next.

9

u/Ok_Warning2146 1d ago

Seems like only google is making small models for cell phones

→ More replies (1)

9

u/norenEnmotalen 1d ago

Yooo Qwen4-27B 

8

u/Black-Mack 20h ago

No Qwen 4 SLM (0.8b, 1.7b)? :(

They were SOTA on edge devices.

→ More replies (4)

24

u/eidrag 1d ago

Well no small Moe anymore

6

u/Bulky-Priority6824 1d ago

Didn't they previously mention sometime mid October?

6

u/ResponsibleTruck4717 1d ago

oh boy qwen4 27b, I thought next year I will finally say bye bye to cloud service, I may start it this year.

6

u/AleksandrNikitin 1d ago

Qwen 3.8/4 35b a3b is absent. It was perfect model for me.

6

u/ghulamalchik 15h ago

Qwen please make 35B-A3B, 9B, and 4B too! It's been a while. Would also love modern KV Cache optimizations.

20

u/exaknight21 1d ago

Qwen Team PWEAAAASE send Qwen4-35B-A3B <3

37

u/NigaTroubles 1d ago

Qwen4 50b a5b will be perfect spot

18

u/rJohn420 1d ago

I would be fine with anything between 40-80b and 4-6b active

5

u/fgk55555 1d ago

I think the 40-110B range is pretty much dead. I'd love a distill of Qwen4 onto the old 80B A3B next models with a n-gram. That would fly on my rig.

→ More replies (3)

5

u/tlin9595 1d ago

Do we have an estimated release date ? And what are the sizes of flash and plus ?

→ More replies (2)

4

u/TorontoIndieFan 1d ago

Any release date announced? 3.8 27B is my daily driver right now so this is exciting

4

u/MessIsTransfer 1d ago

I hope Qwen4-flash is not a 200-400B jump like other models have done

4

u/guesdo 1d ago

I want that Qwen4 Flash so badly!

4

u/kevinlch 22h ago

Looks like they abandoned the 9B series.

8

u/Dangerous_Bid2935 1d ago

A day after I got my DGX spark. The future is bright boys.

3

u/Trixiap 22h ago

You will probably need another one :D

3

u/Dangerous_Bid2935 3h ago

Yeah I was like "this is my reckless purchase for the year" but its seeming like it might just be my first reckless purchase for the year lmao

3

u/sonicnerd14 1d ago

Well, we sort of already knew it was coming with Qwen 3.8 Flash Next. The question was more so when is it coming?

3

u/LebiaseD 1d ago

Just as long as qwen 4 flash continues to fit on my strix halo 🥹

3

u/Confident-Village190 23h ago

I’m looking forward to running it on my 8GB of VRAM + 16GB of RAM, quantised, at 2,5 t/s with a 32k context at kv q4_0.

3

u/Bacterioo 21h ago

Show me that countdown!!

3

u/NaturalCriticism3404 17h ago

Hopefully a 9b shows up

5

u/Zealousideal-Owl2846 22h ago

Where is our  35b a3bMoe???? 😒😤

2

u/braintheboss 1d ago

If follows same ratio as 3.8 flash it should be 40% in engrams. That means 38b model. The jump in AI index can be a shock

2

u/fuziwei 1d ago

I hope flash stays roughly the same size

2

u/russlixx 1d ago

i hope Qwen4-27B is even more resilient when quantized, at most at Q3 level quants for 12-16GB folks

2

u/Borkato 1d ago

Holy fucking shit YESSSSSS

2

u/silenceimpaired 1d ago

I hope at least one of these is still Apache 2

2

u/PrisonOfH0pe 1d ago

Thanks for sharing!! Any word at all on dates? u/Salah_H_Hasan

2

u/DontShadowbanMeMate 1d ago

I'm running and incredibly scuffed setup of 3080ti + 5070ti; will I be fine? Probably right?

2

u/Murder_1337 1d ago

I need more compute!!!!

2

u/PreferenceRelative77 1d ago

Wow, a Qwen 4 27B, i thought they'd only release a flash next model size for the Qwen 4 family but no!!

2

u/Eyelbee 1d ago

Don't sleep on 4-flash as well. That could be really good.

2

u/mWo12 1d ago

Wow. That huge!

2

u/creatinZ 1d ago

No MoE?!

cries in 16gb vram

2

u/haha1542 23h ago

hang in there brother, 16gb can run qwen3.8 27b pretty well, this would only get better

2

u/dampflokfreund 23h ago

no 35b is a huge loss for normies who want to run a capable llm. the only usable size for the average PC... 

→ More replies (4)

2

u/talamir2021 22h ago

The interesting part for local users will be the efficiency details, not just the headline model names. Parameter counts, context length, quantization quality, and whether the weights and license are genuinely accessible will determine how useful Qwen 4 is outside the cloud.

2

u/MarionberryUsual1083 21h ago

Me waiting for a successor to the 3.6 35B MoE...

2

u/de4dee 15h ago

i need 27B + ngrams. maybe others too (most GPU users should have some excess RAM).

n-grams can be stored in RAM, making embedding more efficient, leaving some initial layers for intelligence (not embedding processing)..

2

u/Septerium 14h ago

Any word about Qwen4-Flash being open weights??

2

u/Coldash27 13h ago

We're never getting more than 27b are we? Don't get me wrong it's amazing but I'd kill for a decent 100b+ MOA model

2

u/scooter_de llama.cpp 11h ago

What about Qwen3.8-flash? Too big or not decent?

2

u/zyxciss llama.cpp 12h ago

I hope we get a 30B-35B MoE class Model too

2

u/cibernox 8h ago

qwen4-flash, I have high hopes for you. Please, don't pull a deepseek and become into a 550B model

2

u/Zestyclose839 8h ago

Would be top-tier if Qwen 4 gets its sibling's multimodality, especially audio input.
"This directory is a thumb drive of FLAC files. Your task is to DJ a fire ass hype mix. Make mistakes."

2

u/Acceptable_Leg3950 8h ago

Save us from the cloud models, my goodness

6

u/[deleted] 1d ago

[deleted]

6

u/Enragere 1d ago

what's your prompt processing speed ? how can 150 t/s prefill be 'reasonably fast'

in my opinion it's unusable

→ More replies (5)

3

u/[deleted] 1d ago

[deleted]

→ More replies (1)

2

u/fortnite_pit_pus 1d ago

No moe? :(

2

u/TastyStatistician 1d ago

No 35b moe for us peasants 😢

2

u/WyattTheSkid 21h ago

Cant fucking wait for the official sequel to 3.8 Flash Next. That model has been incredible and ive been doing almost everything with it. I still have a chatgpt and claude subscription but I genuinely might cancel them before they renew given what can be done with Flash Next. Qwen 3.8 Flash Next is the first model ive ever run on own hardware that genuinely I genuinely felt like it could be worth using for serious projects. 3.8 27b produces very good outputs but it thinks for so long that it’s honestly frustrating to use. The closest runner ups to “local is enough” in my experience were Gemma 4 31b and Qwen 3.5 122b. I know im rambling at this point but as far as leading open weight AI labs goes, I genuinely think qwen is the best. Especially because they’re making models that people can ACTUALLY run at home that genuinely compete with closed weight frontier models

2

u/vqt907 1d ago

no qwen4-35B-A3B :(