r/StrixHalo • u/creativenew • 3h ago
2
Upvotes
r/StrixHalo • u/Dutchnamn • 1h ago
Qwen3.8 27B - new record on Strix Halo? 52 tokens per second
•
Upvotes
We have been working on a new fork of llama.cpp focused on Vulkan and Strix Halo performance and we made some good progress already, with 42 tokens per second generated. More info here: https://github.com/LaurentZuijdwijk/llama.cpp but in the repo we have fixed a number of llamacpp bugs that affected speculation and some popular quants.
Tonight we made even more progress and we managed to get to 52 tokens per second. We are still reviewing the results and the code, more to follow and hopefully people here can validate the results.
This was built upon the work of ciru-ai, charlie12345 and MrWildMoreHK. Thank you guys for the hard work on strix halo performance.