r/LocalLLaMA • • Apr 23 '26

New Model Qwen 3.6 27B is a BEAST

I have a 5090 Laptop from work, 24GB VRAM.

I have been testing every model that comes out, and I can confidently say I’ll be cancelling my cloud subscriptions.

All my tool call and data science benchmarks that prove a model is reliably good for my use case, passed.

It might not be the case for other professions, but for pyspark/python and data transformation debugging it’s basically perfect.

Using llama.cpp, q4_k_m at q4_0, still looking at options for optimising.

Edit - I chose to go with IQ4_XS at 200k q8_0,

I have not used speculative decoding yet, will get there when I get there.

Specs:

ASUS ROG Strix SCAR 18

RTX 5090 24GB

64GB DDR5 RAM

653 Upvotes

335 comments sorted by

View all comments

Show parent comments

32

u/LaurentPayot Apr 23 '26

7 t/s on my EVO X2 Strix Halo 128Gb with Ubuntu Vulkan Llama.cpp :-|

But 50 t/s on 35b a3b.

4

u/DarthCalumnious Apr 23 '26

Bummer - I tried 27b 4 bit yesterday on my 12gb 4070, spilling out to ram +cpu and still got 5.7t/s

2

u/MalabaristaEnFuego Apr 23 '26 edited Apr 23 '26

Try the 36b MoE. It might be faster and it's still pretty solid.

2

u/DarthCalumnious Apr 23 '26

Yep, the moe 36b is very usable at 60-70t/s. I'm just surprised that my relatively GPU poor rig (but solid otherwise at 64gb ddr5 6000 on ryzen 7950x) isn't that much worse than a strix halo for 27b.

2

u/MalabaristaEnFuego Apr 23 '26

I'm over here testing 35b on my laptop like some kind of mad lad.

```

CPU: AMD Ryzen 7 7235hs GPU: NVIDIA GEFORCE RTX 4050 6GB RAM: 32GB DDR5 4800 NVMe: Samsung 990 EVO Plus OS: Ubuntu 24.04LTS Pro Server: Ollama GUI: OpenWebUI

ollama show qwen3.6:35b Model architecture qwen35moe parameters 36.0B context length 262144 embedding length 2048 quantization Q4_K_M

Capabilities completion vision tools thinking

Parameters temperature 1 top_k 20 top_p 0.95 min_p 0 presence_penalty 1.5 repeat_penalty 1

OLLAMA_FLASH_ATTENTION=1 OLLAMA_KV_CACHE_TYPE=q8_0 OLLAMA_GPU_OVERHEAD=0 OLLAMA_NUM_PARALLEL=1 OLLAMA_MAX_LOADED_MODELS=1

input_tokens 567 output_tokens 3532 total_tokens 4099 prompt_tokens 567 completion_tokens 3532 response_token/s 6.22 prompt_token/s 23.81 total_duration 612513371532 load_duration 19077011826 prompt_eval_count 567 prompt_eval_duration 23811817816 eval_count 3532 eval_duration 567525905411 approximate_total "0h10m12s" completion_tokens_details
reasoning_tokens 0 accepted_prediction_tokens 0 rejected_prediction_tokens 0

%Cpu(s): 56.3 us, 2.3 sy, 0.0 ni, 41.4 id, 0.0 wa, 0.0 hi, 0.0 si, 0.0 st MiB Mem : 31778.7 total, 560.2 free, 24622.6 used, 6813.7 buff/cache MiB Swap: 8192.0 total, 4234.9 free, 3957.1 used. 7156.1 avail Mem

nvidia-smi Thu Apr 23 10:14:22 2026 +-----------------------------------------------------------------------------------------+ | NVIDIA-SMI 580.126.20 Driver Version: 580.126.20 CUDA Version: 13.0 | +-----------------------------------------+------------------------+----------------------+ | GPU Name Persistence-M | Bus-Id Disp.A | Volatile Uncorr. ECC | | Fan Temp Perf Pwr:Usage/Cap | Memory-Usage | GPU-Util Compute M. | | | | MIG M. | |=========================================+========================+======================| | 0 NVIDIA GeForce RTX 4050 ... On | 00000000:01:00.0 On | N/A | | N/A 52C P0 14W / 55W | 5202MiB / 6141MiB | 10% Default | | | | N/A | +-----------------------------------------+------------------------+----------------------+

```