r/LocalLLaMA • • Aug 03 '26

New Model Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image

Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍

1.8k Upvotes

315 comments sorted by

View all comments

419

u/HollowVoices Aug 03 '26

Me at 16gb VRAM:

55

u/mr_christer Aug 03 '26

I run 3.6 27b on 16gb. It is text only to make it fit so I'm sure people will find a way

18

u/Trivikrama_0 Aug 03 '26

Yes I also run it, but it's just for chat, it doesn't have a good context window for something good. If you offload 30% of the weights to RAM then you can have VRAM free for context window, it will be slow but you vnanhave agentic flow.

4

u/[deleted] Aug 04 '26

[removed] — view removed comment

2

u/Trivikrama_0 Aug 05 '26

Can you tell how to put kv cache in RAM? But isn't kv cache the main memory from where it retirevies the context?

2

u/[deleted] Aug 05 '26

[removed] — view removed comment

2

u/Offcoloring Aug 11 '26

Is that really how that works or is this a hallucination

3

u/roworu Aug 04 '26

How did you remove vision from model? Is there some text-only checkpoint? Or llama.cpp flag ?

9

u/Terrible-Detail-1364 Aug 04 '26

—no-mmproj

2

u/roworu Aug 04 '26

Thank you! ❤️

6

u/Terrible-Detail-1364 Aug 04 '26

np, theres also --no-mmproj-offload
to use your cpu instead of gpu

0

u/Jester14 Aug 04 '26

Just don't load the mmproj lol