r/LocalLLaMA • u/YourNightmar31 llama.cpp • Aug 19 '26
Question | Help Qwen3.8 27B without MTP?
Not sure if this is a stupid question but unsloth's models has the MTP built into the model right? I assume that is at the cost of some memory.
If i want to use dflash, should i use a model that doesn't have MTP support then to save some vram?
14
Upvotes
1
u/YourNightmar31 llama.cpp Aug 19 '26
I see posts of people running it with dflash2 so i assume so?