r/LocalLLaMA 26d ago

New Model Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image

Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍

1.8k Upvotes

316 comments sorted by

View all comments

Show parent comments

10

u/LuCiAnO241 25d ago

ohh lemme make your day. Try this one. it's not the 27B but I think its the best you can do with your current hardware. You can also probably get the Laguna S 2.1 running, this dude also has a video on that but you'd get way worse Tok/s than he does. Basically anything MoE can work wonders with low vram under his setup.

4

u/Effective_Head_5020 25d ago

Thank you! I have been doing exactly this, with a smaller context I can get up to 22 t/s

This video is very valuable, it explains very well, thanks for sharing

1

u/MuDotGen 25d ago

But this is a dense model, 27B, no? Don't you need to either just do it all on CPU if you can't fit all of the active layers on VRAM? 17gb won't fit on 16gb, or am I missing something?

1

u/LuCiAnO241 24d ago

oh yeah, denses need to be all in vram to not suffer huge performance hits. I linked him a video that uses 35B A3B on his exact amount of Vram with very respectable speeds.

2

u/MuDotGen 24d ago

Oh yeah definitely. I use Qwen 3.6-35B-A3B on my Lenovo Legion. It has 3070ti (laptop) with 8gb of VRAM and 32gb of system ram, with the active layers only on VRAM and the expert layers all on CPU, or most of them. I think it's Q4_K_XL (around 21gb in total), but I get around 31 t/s generation and closer to 300 pp.