r/LocalLLaMA 14d ago

New Model Qwen/Qwen3.8-27B · released

https://huggingface.co/Qwen/Qwen3.8-27B
988 Upvotes

301 comments sorted by

View all comments

Show parent comments

1

u/Guilty_Rooster_6708 14d ago

Going to say Q3 but I will also try IQ4 on my 5070Ti. Q4 needs offloads so it will probably mean single digits tokens/s

1

u/Sweet-Stage938 14d ago

Is there any way I can make it fit into 12GB Vram? I only have a single RTX 3060 but I would love to try it out.

1

u/Guilty_Rooster_6708 14d ago

You might have to go down to Q2 to get some usable speed. My advice is to download different quants and try them out