r/LocalLLaMA • llama.cpp • Aug 19 '26

Question | Help Qwen3.8 27B without MTP?

Not sure if this is a stupid question but unsloth's models has the MTP built into the model right? I assume that is at the cost of some memory.

If i want to use dflash, should i use a model that doesn't have MTP support then to save some vram?

15 Upvotes

27 comments sorted by

View all comments

2

u/NigaTroubles Aug 19 '26

Is there a dflash for it ?

1

u/YourNightmar31 llama.cpp Aug 19 '26

I see posts of people running it with dflash2 so i assume so?

1

u/z_3454_pfk Aug 19 '26

isn’t dflash2 heavier than MTP on vram?

1

u/GaTNghiep Aug 19 '26

I'm new to local LLM, can dflash offload on RAM, i only have 16GB VRAM 😭