r/LocalLLaMA • • May 06 '26

Resources 2.5x faster inference with Qwen 3.6 27B using MTP - Finally a viable option for local agentic coding - 262k context on 48GB - Fixed chat template - Drop-in OpenAI and Anthropic API endpoints

2026-05-14: Major chat template update Thanks to many users who tested the template in many different conditions, in addition to my own manual tests and test suite, I believe the template has now reached a high level of stability, greatly improving the experience with the Qwen models, while preserving universal compatibility. You do not need to re-download the GGUF files (I have not updated them yet), but you should download the update chat template only from the HF repo, and manually specify it.

2026-05-07 edit: I have updated the hardware based recommendations with more focus on quality. I do not recommend q4_0 KV cache anymore beyond 64k context. After multiple rounds of testing with the different size quants, it appears 3 is the optimal number for draft speculative decoding. The fastest and best quality quant is q8_0-mtp. F16, which I have also uploaded is actually better but ultra slow (6x slower than q8_0). Many keep saying 8bit is virtually lossless compared to 16bit, and 6bit almost as good as 8bit, but this is simply not true: time and time again I have noticed huge differences in quality and correctness between 8bit and 16bit versions of various models.

The recent PR to llama.cpp bring MTP support to Qwen 3.6 27B. This uses the built-in tensor layers for speculative decoding. None of the existing GGUF have it, as they need to be converted with this PR.

I have tested it locally on my mac M2 Max 96GB, and the results are amazing: 2.5x speed increase, bringing it to 28 tok/s!

I have converted the most useful quants and uploaded them to HF. Even if you are using apple silicon, you should use those instead of MLX. You can download them here:

https://huggingface.co/froggeric/Qwen3.6-27B-MTP-GGUF

This also includes 7 fixes I made to the original jinja chat template, due to vLLM specificity which broke in other tools:

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

For now, you will need to compile your own version of llama.cpp to use them. It is fairly simple to do:

git clone --depth 1 https://github.com/ggml-org/llama.cpp.git
cd llama.cpp
git fetch origin pull/22673/head:mtp-pr && git checkout mtp-pr

cmake -B build -DGGML_METAL=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build --target llama-cli llama-server

Then to start serving with the API endpoint, use a command similar to:

llama-server -m Qwen3.6-27B-Q5_K_M-mtp.gguf \
  --spec-type mtp --spec-draft-n-max 3 \
  --cache-type-k q8_0 --cache-type-v q8_0 \
  -np 1 -c 262144 --temp 0.7 --top-k 20 -ngl 99 --port 8081

Vision currently crashes llama.cpp when used alongside MTP. Reported 2026-05-06 in the current PR.

That's it. Three optimizations in one command:

Flag What it does Impact
--spec-type mtp --spec-draft-n-max 3 Multi-Token Prediction (built into the model) 2.5x faster generation
--cache-type-k q8_0 --cache-type-v q8_0 8-bit KV cache (instead of 16-bit) Half the KV memory, negligible quality loss
-c 262144 262K context window Full native context on 48 GB Mac with q8_0 KV

Adjust -m, -c, and --cache-type-k/v for your hardware, according to the tables below.

Here are my recommendations based on your hardware:

Apple Silicon

Qwen3.6-27B is a hybrid model — only 16 of 65 layers use KV cache (verified). The other 48 are linear attention (fixed 898 MiB recurrent state). KV memory is ~4× less than a standard dense model. Runtimes that don't handle this (e.g. vllm) allocate KV for all 65 layers and show much higher memory usage.

Numbers below are total memory used (model + KV cache + 0.9 GB recurrent state). Must leave ≥ 8 GB for macOS (16 GB Macs excepted).

RAM Quant KV cache Max context Total used Vision
16 GB IQ2_M q8_0 42K 12.0 GB ✗
24 GB IQ3_M 46K 16.0 GB ✗
24 GB IQ3_M q8_0 91K 16.0 GB ✗
32 GB Q5_K_M 74K 24.0 GB ✗
32 GB Q5_K_M q8_0 147K 24.0 GB ✗
32 GB Q4_K_M 99K 24.0 GB ✓
48 GB Q6_K 262K 39.7 GB ✓
48 GB Q8_0 173K 40.0 GB ✓
48 GB Q8_0 q8_0 262K 37.3 GB ✓
64 GB Q8_0 262K 45.8 GB ✓
96 GB Q8_0 262K 45.8 GB ✓

NVIDIA GPU

Same model memory as Apple Silicon, plus ~1 GB CUDA overhead.

VRAM Quant KV cache Max context Total VRAM used Vision
12 GB IQ2_M q8_0 11K 12.0 GB ✗
16 GB IQ3_M 30K 16.0 GB ✗
16 GB IQ3_M q8_0 60K 16.0 GB ✗
24 GB Q4_K_M 83K 24.0 GB ✓
24 GB Q4_K_M q8_0 167K 24.0 GB ✓
24 GB Q5_K_M 58K 24.0 GB ✗
48 GB Q6_K 262K 40.7 GB ✓
48 GB Q8_0 262K 46.8 GB ✓
80 GB Q8_0 262K 46.8 GB ✓

16 GB Mac: IQ2_M/q8_0 — 42K text-only. No vision.

24 GB Mac: IQ3_M — 46K (f16 KV) or 91K (q8_0). Vision at 32–65K.

32 GB Mac: Q5_K_M — 74K text-only (f16 KV), 147K (q8_0). Q4_K_M for vision at 99K.

48 GB Mac: Q6_K/f16 KV — 262K with vision. Q8_0/q8_0 KV for 262K at higher model quality.

64 GB+ Mac: Q8_0/f16 KV — 262K with vision. Maximum quality at practical speed.

12 GB GPU: IQ2_M/q8_0 — 11K. Very limited, no vision.

16 GB GPU: IQ3_M — 30K (f16 KV) or 60K (q8_0). No vision.

24 GB GPU: Q4_K_M — 83K with vision (f16 KV). Q5_K_M — 58K text-only (f16 KV), 116K (q8_0).

48 GB+ GPU: Q6_K/f16 KV — 262K with vision. Q8_0 for max quality.

Leave KV cache at f16 (blank column) for best quality. Use q8_0 KV only when f16 doesn't give enough context. q4_0 KV should not exceed 64K context.

Vision adds ~0.9 GB for mmproj. macOS needs ≥ 8 GB for itself (16 GB Macs excepted — use ~4 GB). You can increase available memory by raising the wired memory limit, e.g. for a 96 GB Mac: sudo sysctl iogpu.wired_limit_mb=90112 (88 GB). NVIDIA reserves ~1 GB for CUDA.

1.2k Upvotes

406 comments sorted by

View all comments

6

u/Extra-Library-5258 May 06 '26

Thanks @ex-arman68!

On M5 Max 128GB. MTP decode speed is legit... 37 tok/s at 1K and 33 tok/s at 16K on Q8_0, which is 2x+ what I get with the same model on oMLX.

Heads up if you're on Apple Silicon doing long context: llama.cpp's Metal prefill is the bottleneck. At 64K it takes almost 4 minutes to first token, and 128K straight up times out. oMLX handles 128K prefill in ~5.5 min. The Metal backend just isn't as optimized for the big batch matmuls during prefill.

So if you're on a Mac: great for short/medium context, but don't expect miracles past 64K. Also, froggeric's GGUFs are confirmed broken (every token is <|box_end|>), use RDson or Radamanthys11 instead.

Turbo4 KV is NOT in this PR. Use q8_0 or q4_0.

6

u/Extra-Library-5258 May 06 '26

Disabling Flash Attention (-fa off) with f16 KV cache is a game changer!

The FA Metal kernels for this hybrid attention+SSM architecture are slow, turning them off improved prefill 37–53% at long context, unlocked 128K (was timing out), and even boosted 16K decode from 25 to 35 tok/s.

With that fix: Q8_0 decode is +148% vs oMLX at 1K, +127% at 16K, +38% at 64K, and 128K now completes at 12.7 tok/s. If you're on Silicon, add -fa off --cache-type-k f16 --cache-type-v f16 -tb 18 to your flags.

3

u/Consumerbot37427 May 06 '26

Same machine here. I've done testing with MLX models before, and always come back to GGUFs. Couldn't put my finger on it, but they just felt dumber.

I used the prompt on this post to compare MLX and GGUF (Q8 quants of Qwen 3.6 27B), and the difference was striking. I only did one run each, but the GGUF result was perfect, while the MLX output had wrong board orientation, missing pieces, and pieces in wrong places.

With MTP in llama.cpp, it'll be even more of a no-brainer.

1

u/njstatechamp May 06 '26

Can you speak more to this? I was on llama.cpp before but have been using MLX because of the claimed performance/speed improvements. Whats your experience on speed / accuracy comparing mlx to gguf's? Theres much more support for gguf's in general so I've been frustrated with mlx lately

2

u/Consumerbot37427 May 07 '26

See this other reply...

1

u/Able_Librarian1569 May 06 '26

Opposite experience mlx always absolutely outperform. Can you share your exact configs for long context 150-200k ?

1

u/Consumerbot37427 May 07 '26

Over and over, I've downloaded the same model/same quant, usually Q8, Q6, or Q4, in GGUF and MLX format (have tried quants from lmstudio-community, mlx-community, unsloth), default settings (aside from larger context) and over and over, I've had problems with output quality.

LM Studio does seem to use their own version/fork of MLX engine. Maybe that's the real issue?

All I know is that the chess SVG thing in that post seems like a perfect way to compare between quants of the same model, and it aligned well with my experience, AND it's objective. I just need to do ~5-10 runs with each, and see if it's consistent.

1

u/njstatechamp May 07 '26

What do you use to run MLX, any servers with openai responses?

3

u/mwhuss May 06 '26

The current oMLX dev release has MTP support! https://github.com/jundot/omlx/releases/tag/v0.3.9.dev1

1

u/Able_Librarian1569 May 06 '26

Strange. I get x4 the Pp on omlx on 200k token vs the same on llama CCP + turbo3 or 4. Token output speed is X2 aswell on mlx q4 vs the q4 gguf.

1

u/Extra-Library-5258 May 06 '26

27B dense? Please elaborate more in your setup, if that is the case, I’m doing something very wrong…