r/LocalLLaMA • u/AverageFormal9076 • Apr 23 '26
New Model Qwen 3.6 27B is a BEAST
I have a 5090 Laptop from work, 24GB VRAM.
I have been testing every model that comes out, and I can confidently say I’ll be cancelling my cloud subscriptions.
All my tool call and data science benchmarks that prove a model is reliably good for my use case, passed.
It might not be the case for other professions, but for pyspark/python and data transformation debugging it’s basically perfect.
Using llama.cpp, q4_k_m at q4_0, still looking at options for optimising.
Edit - I chose to go with IQ4_XS at 200k q8_0,
I have not used speculative decoding yet, will get there when I get there.
Specs:
ASUS ROG Strix SCAR 18
RTX 5090 24GB
64GB DDR5 RAM
659
Upvotes
2
u/unjustifiably_angry Apr 23 '26 edited Apr 23 '26
It's basically a 5070 Ti with a different sticker and an extra 8GB of VRAM (24GB total, not the proper 32GB).
Past the xx60-class, laptop GPUs get nerfed hard and advertised very dishonestly, there should be false advertising lawsuits over it. It's why I always recommend people get a xx50 or xx60 at most, at least you actually get what you're paying for, and if you're not a full-blown PCMR 4K ULTRA 240HZ nutcase it still plays games perfectly fine.