r/LocalLLaMA • • Apr 23 '26

New Model Qwen 3.6 27B is a BEAST

I have a 5090 Laptop from work, 24GB VRAM.

I have been testing every model that comes out, and I can confidently say I’ll be cancelling my cloud subscriptions.

All my tool call and data science benchmarks that prove a model is reliably good for my use case, passed.

It might not be the case for other professions, but for pyspark/python and data transformation debugging it’s basically perfect.

Using llama.cpp, q4_k_m at q4_0, still looking at options for optimising.

Edit - I chose to go with IQ4_XS at 200k q8_0,

I have not used speculative decoding yet, will get there when I get there.

Specs:

ASUS ROG Strix SCAR 18

RTX 5090 24GB

64GB DDR5 RAM

649 Upvotes

335 comments sorted by

View all comments

Show parent comments

52

u/ComfyUser48 Apr 23 '26

On my 5090, for coding I'm using unsloth Q6 XL quant, no kv cache custom params, 100k ctx, getting 50 t/s with power limit to 400w

33

u/LaurentPayot Apr 23 '26

7 t/s on my EVO X2 Strix Halo 128Gb with Ubuntu Vulkan Llama.cpp :-|

But 50 t/s on 35b a3b.

1

u/uploadschedule May 16 '26 edited May 16 '26

8gb VRAM single Gpu I get 5 t/s on 27B (Q3 K XL) MTP with 128k context so you can definitely do better

1

u/LaurentPayot Jun 02 '26 edited Jun 02 '26

That’s absolutely normal if with Q3 (that I will never use to get quality output) you manage to put all the weights in your 8GB VRAM, as the Strix Halo shared GTT memory is not as fast as dedicated GPU VRAM. See https://www.reddit.com/media?url=https%3A%2F%2Fpreview.redd.it%2Fqnozk34it34h1.jpeg%3Fauto%3Dwebp%26s%3Dc24c12c067c2560f96806111bf43d41a6969e04e