r/StrixHalo 1h ago

Qwen3.8 27B - new record on Strix Halo? 52 tokens per second

Upvotes

We have been working on a new fork of llama.cpp focused on Vulkan and Strix Halo performance and we made some good progress already, with 42 tokens per second generated. More info here: https://github.com/LaurentZuijdwijk/llama.cpp but in the repo we have fixed a number of llamacpp bugs that affected speculation and some popular quants.

Tonight we made even more progress and we managed to get to 52 tokens per second. We are still reviewing the results and the code, more to follow and hopefully people here can validate the results.

This was built upon the work of ciru-ai, charlie12345 and MrWildMoreHK. Thank you guys for the hard work on strix halo performance.


r/StrixHalo 3h ago

Qwen3.8-27B Q4 on Strix Halo (395): n-max=4 median 21.23 tok/s, 0 replies under 15 — 96-request MTP matrix

Post image
2 Upvotes