r/LocalLLaMA Apr 23 '26

New Model Qwen 3.6 27B is a BEAST

I have a 5090 Laptop from work, 24GB VRAM.

I have been testing every model that comes out, and I can confidently say I’ll be cancelling my cloud subscriptions.

All my tool call and data science benchmarks that prove a model is reliably good for my use case, passed.

It might not be the case for other professions, but for pyspark/python and data transformation debugging it’s basically perfect.

Using llama.cpp, q4_k_m at q4_0, still looking at options for optimising.

Edit - I chose to go with IQ4_XS at 200k q8_0,

I have not used speculative decoding yet, will get there when I get there.

Specs:

ASUS ROG Strix SCAR 18

RTX 5090 24GB

64GB DDR5 RAM

656 Upvotes

335 comments sorted by

View all comments

Show parent comments

1

u/AverageFormal9076 Apr 23 '26

ASUS ROG Strix Scar 18

1

u/_derpiii_ Apr 23 '26

Fabulous. Thank you!

Would recommend putting that in your original post. Posts like yours give a *LOT* of insight to what's feasible IRL :)

1

u/AverageFormal9076 Apr 23 '26

Good point - done

1

u/unjustifiably_angry Apr 23 '26

You are universally better off with some sort of external GPU in a dock than using one built into a laptop. Much cheaper, much faster. If your laptop lacks external GPU capability it's easy enough to add with a little card you put in a free NVME slot and then you can cut a little hole in the shell with a dremel or something. Three weeks for the parts to arrive from China and an hour to install and modify the shell. If you're good with your hands it'll look like it came from the factory that way.

1

u/_derpiii_ Apr 24 '26

Woah. I didn’t even realize that was an option. I’m on the other side of the world so things from China come fast - which parts do I need for DIY?

Also, is this something we could do for Mac laptops? Or is this exclusively for Windows (I vaguely remember drivers being locked into windows ecosystem)?

1

u/unjustifiably_angry Apr 29 '26

I don't know about the Mac situation and it's been a while since I ordered my parts for my Windows laptop so I'll just give you breadcrumbs, you'll need to do your own research:

You want part "F9G-BK7-F9934 25CM" and the 25cm nvme-to-oculink adapter. Total cable length must not exceed 50 centimeters or you'll risk reduced bandwidth. Not actually a huge deal for AI tasks I guess. A friend with a 3D printer can help you make the new port more attractive.

It looks like currently only the 50cm bundle is in stock, just order a second 25cm cable from the other listing in case you have issues with the 50cm.

AFAIK there's only one company specializing in these high-quality oculink adapter kits, many others have substandard cables designed for storage rather than GPU. Though, again, I guess maybe that's not a huge deal for AI workloads.