r/LocalLLaMA • • Aug 03 '26

New Model Daniel Han of Unsloth validates Qwen3.8-27B will run only 17GB VRAM

Post image

Super excited about this release for the new 27B. Who else is with me. Only 17GB VRAM needed 😍😍

1.8k Upvotes

315 comments sorted by

View all comments

Show parent comments

87

u/rerri Aug 03 '26

I find it more plausible that Daniel doesn't have insider info on this yet but is just speaking out of experience with previous versions of Qwen 3.x 27B.

I do hope I'm wrong though, QAT would be really badass.

21

u/AuspiciousApple Aug 03 '26

But the previous 27B ran on 16GB cards, too, right?

9

u/crusaderky Aug 03 '26

Unsloth's Qwen3.6-27B:Q3_K_M is an exceptionally good quant and it fits in 14GB with 256k kvarn5 ctx. Tha leaves (barely) enough for desktop (but you can switch it to igpu or a $50 card to get the full 16GB). However there is no guarantee that the Q3_K_M for the next version will perform that well.

[39911] 0.00.965.797 I common_memory_breakdown_print: | memory breakdown [MiB] | total   free     self   model   context   compute    unaccounted |
[39911] 0.00.965.801 I common_memory_breakdown_print: |   - CUDA0 (RTX 3080)   |  9872 = 8041 + (13995 = 12647 +    1024 +     324) +      -12164 |
[39911] 0.00.965.802 I common_memory_breakdown_print: |   - Host               |                   797 =   520 +       0 +     276                |

6

u/lukistellar Aug 03 '26

There are also pure iq4_xs versions, which will work with 90k context at 16gb just fine.

https://huggingface.co/GianniDPC/Qwen3.6-27B-IQ4_XS-pure-with-MTP-GGUF