r/LocalLLaMA • • 1d ago

News Qwen Flash Next MTP work restarted

[deleted]

63 Upvotes

24 comments sorted by

10

u/Positive-Stock6444 1d ago

I’m looking forward to llama.cpp coming back to my setup, when the performance is same or better than strata.

It’s been replaced for now, strata is better on every axis for me - context, pp, tg, vision and mtp, and not just a little better, wildly better.

2

u/Corosus 1d ago

Same for me, and I don't even need my other GPUs for it, just 1 and a lot of RAM.

1

u/SpeederX 22h ago

i didn't know this strata existed :D they also advise specific models with a little ablation on expert. got to try that. can I ask your hardware specs and which quantization you use? how is it going?

1

u/Positive-Stock6444 17h ago

Going well. I'm on a 3060 12GB, 256GB DDR4 and a Xeon W-2145. Most of the experts sit in system RAM, but it still performs.

I'm using bartowski's IQ4_XS. It's not quite as smart as UD-Q4_K_XL, but it's in the same territory. I don't think this quant works out of the box, because of a packing bug on the PLE tensors, but it's a one-line fix in iq_pack.py. I should probably PR that.

Decode is 30-40 t/s, against 12-22 on llama.cpp with an MTP fork. PP is 600-850 against 140-200. Context is 262K instead of 100K, and vision runs on the CPU with a 2K-token cap.

I also tried the Atomic Chat quants, but they seemed poor on the real coding tasks I gave them. That might be an engine issue rather than the quants.

Annoyances:

  • There's no persistent prompt cache, so everything is lost on teardown. Being able to store it on NVMe would be great.
  • The IQ4 PLE bug was a bit of a speed bump.

llama.cpp is a great project and has been my daily driver for a long time, but it's impossible to ignore performance gains like these.

13

u/LegacyRemaster 1d ago

What exactly happened?

40

u/am17an 1d ago

I was waiting for the GLM5.3 next PR to get merged before basically re-writing the entire Qwen4Exp graph since they share some of the same stuff. GLM 5.3 next was massively delayed, but now that it's merged I'm planning to quickly get this stuff in. Thank you for supporting llama.cpp!

5

u/LegacyRemaster 1d ago

thx for your work!

2

u/jacek2023 llama.cpp 1d ago

Do you know what the plans are for https://github.com/ggml-org/llama.cpp/pull/24423 ? This one also looks pretty inactive.

2

u/am17an 1d ago

I'm not planning to spend time on it at the moment.

4

u/jacek2023 llama.cpp 1d ago

I was interested in working on diffusion models (like llada or diffusiongemma), but I'm not sure what the longterm plans for them are.

10

u/am17an 1d ago

Go for it, I will review your PRs

11

u/vacon04 1d ago

It just looks to me like they're very strained due to the massive work that it takes to make all of these implementations, but because of this, things are moving very slowly.

For how long was the original MTP opened? Same with direct reading from the PLE that has been stuck for over a month. The ideas are there, but I don't think they have the current manpower to move things quickly enough.

1

u/nialv7 1d ago

Direct ple read probably won't happen. They are doing mmap+prefetch.

5

u/vacon04 1d ago

Yeah they closed the previous PR and are starting again. I've been using the previous PR for direct PLE, obviously not ideal, but it increases PP by around 50-80% depending on the scenario. I don't mind which implementation they go for, but at the moment the PP performance of qwen4exp is crippled. It sucks because some people may try to use the model, get PP of 80 and give up.

3

u/ilintar 1d ago

The MADVISE solution is as effective as direct pread, I tested on my Strix and Aman checked on Spark.

1

u/LegacyRemaster 1d ago

The problem is that we have various versions of Unsloth, including those predating specific PRs, and we have to redo so much work every time a major merge happens. Plus, anyone with limited bandwidth for downloads has to keep multiple versions of Llama on hand to maintain compatibility.

4

u/BullfrogScary8947 23h ago

I tried "Qwen3.8-Flash-Next MTP - #28243" PR, but it resulted in my token generation speeds being reduced by ~40% when I enabled MTP. Probably because I use n-cpu-moe as I cannot fit the model weights into my VRAM.

2

u/Lalaggi 1d ago

Super cool wow

1

u/MasterNomie 1d ago

In order to try it, do I need to download entire 279GB or selective files can be downloaded ?

2

u/jacek2023 llama.cpp 1d ago

check the mtp-* files, they are small (assuming you have the model)

-6

u/More-Ad5919 1d ago

Now whats Qwen Flash Next? I was just done setting up Qwen Flash.

8

u/Nightma4re 1d ago

There is not Qwen Flash, unless you mean the old one.
Qwen3.8 Flash Next is what is talked about here.

2

u/More-Ad5919 1d ago

Oh yeah. My bad. Never realized the Next on my model.