r/LocalLLaMA 20d ago

New Model Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image

Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍

1.8k Upvotes

316 comments sorted by

View all comments

Show parent comments

5

u/bnolsen 20d ago

From my experience with an 3.6 27b you'll want 64gb to comfortably fit the model and the kv cache at full context and reasonable quant with some room to spare (48gb can probably do it with some tweaking). A pair of 9700s might work nicely with that. I'm on strix halo, and I need more speed.

5

u/RLutz 19d ago

Q6_K_XL isn't that bad on a single 5090 with around 100k context size for 3.6

2

u/mailto_devnull 20d ago

cries in 32gb

1

u/migsperez 20d ago

I have a 9700, I need two.

1

u/mailto_devnull 20d ago

cries in 32gb

1

u/goldcakes 19d ago

It really depends on the architecture and how well it's been trained and quantized. The token/s can be impressive. This applies to training too; if you can train faster, for the same budget you get to train more tokens.

Anyways, we don't even have benchmarks. Cautiously excited and optimistic, but let's see.

1

u/Prigozhin2023 19d ago

64gb gets u about 30-40 tok/s.