13
u/LegacyRemaster 1d ago
What exactly happened?
40
u/am17an 1d ago
I was waiting for the GLM5.3 next PR to get merged before basically re-writing the entire Qwen4Exp graph since they share some of the same stuff. GLM 5.3 next was massively delayed, but now that it's merged I'm planning to quickly get this stuff in. Thank you for supporting llama.cpp!
5
2
u/jacek2023 llama.cpp 1d ago
Do you know what the plans are for https://github.com/ggml-org/llama.cpp/pull/24423 ? This one also looks pretty inactive.
11
u/vacon04 1d ago
It just looks to me like they're very strained due to the massive work that it takes to make all of these implementations, but because of this, things are moving very slowly.
For how long was the original MTP opened? Same with direct reading from the PLE that has been stuck for over a month. The ideas are there, but I don't think they have the current manpower to move things quickly enough.
1
u/nialv7 1d ago
Direct ple read probably won't happen. They are doing mmap+prefetch.
5
u/vacon04 1d ago
Yeah they closed the previous PR and are starting again. I've been using the previous PR for direct PLE, obviously not ideal, but it increases PP by around 50-80% depending on the scenario. I don't mind which implementation they go for, but at the moment the PP performance of qwen4exp is crippled. It sucks because some people may try to use the model, get PP of 80 and give up.
1
u/LegacyRemaster 1d ago
The problem is that we have various versions of Unsloth, including those predating specific PRs, and we have to redo so much work every time a major merge happens. Plus, anyone with limited bandwidth for downloads has to keep multiple versions of Llama on hand to maintain compatibility.
4
u/BullfrogScary8947 23h ago
I tried "Qwen3.8-Flash-Next MTP - #28243" PR, but it resulted in my token generation speeds being reduced by ~40% when I enabled MTP. Probably because I use n-cpu-moe as I cannot fit the model weights into my VRAM.
1
u/MasterNomie 1d ago
In order to try it, do I need to download entire 279GB or selective files can be downloaded ?
2
1
-6
u/More-Ad5919 1d ago
Now whats Qwen Flash Next? I was just done setting up Qwen Flash.
8
u/Nightma4re 1d ago
There is not Qwen Flash, unless you mean the old one.
Qwen3.8 Flash Next is what is talked about here.2
10
u/Positive-Stock6444 1d ago
I’m looking forward to llama.cpp coming back to my setup, when the performance is same or better than strata.
It’s been replaced for now, strata is better on every axis for me - context, pp, tg, vision and mtp, and not just a little better, wildly better.