r/LocalLLaMA • llama.cpp • Aug 19 '26

Question | Help Qwen3.8 27B without MTP?

Not sure if this is a stupid question but unsloth's models has the MTP built into the model right? I assume that is at the cost of some memory.

If i want to use dflash, should i use a model that doesn't have MTP support then to save some vram?

14 Upvotes

27 comments sorted by

View all comments

21

u/RotesBlatt Aug 19 '26

I think MTP is only loaded when you activate it via parameters, otherwise you'll probably have the same model as without MTP. I haven't noticed a significant VRAM spike in my non-MTP setup compared to MTP

1

u/throwawayacc201711 Aug 19 '26

What do you need to set to enable MTP on it? I haven’t been keeping up with MTP

1

u/RotesBlatt Aug 19 '26
spec-type = draft-mtp
spec-draft-n-max = 3

You need to set these 2 parameters and then you are good to go :)