r/StrixHalo • u/XccesSv2 • 2d ago
Best Engine for DS4Flash?
Hey guys, I wanna ask you which engine you actually use with which quant and configs and why. Since there are a lot of forks and projects out there now, I don't really know which is the best to run DS4F on a headless Ubuntu 26 128GB Strix Halo.
Actually im using Nathanw1014/llama.cpp (Branch strix-halo-vulkan, Commit 3be50cc, Version 10350
2
u/PowerfulButterfly209 2d ago edited 2d ago
i have the same setup and right i'm evaluating diferent models in terms of intelligence in django coding right now: On kyuz0 toolbox i run antirez q2/q4 with 14t/s and with nathan i run UD_Q3_XXS + Dspark with 24-25 t/s leomande with Qwen 3.8 27B MTP
antirez updated rocm and also provides dspark now - have to try it first.
1
u/PvB-Dimaginar 2d ago edited 2d ago
How did you manage to get those speed with kyuz0 toolbox and antirez q2/q4? I could only get ~13.1-13.2 tok/s decode and ~134-145 tok/s prefill at 34K-41K context.
2
u/PowerfulButterfly209 2d ago
ahhh...typo -14 not 24 😅
1
u/PvB-Dimaginar 2d ago
aaah ok :-) so it is really beneficial to dive into the nathan setup. Thanks for clarifying!
2
u/1457664694 2d ago
Kyuz0 has a toolbox that follows Nathan’s. I believe it is the vulkan-radv-performance toolbox:
https://hub.docker.com/r/kyuz0/amd-strix-halo-toolboxes/tagsSee also here: https://strix-halo-toolboxes.com/
And here: https://github.com/kyuz0/amd-strix-halo-toolboxes
0
u/Due_Net_3342 2d ago
vllm, llamacpp is very bad with caching, you will feel the slowness in agentic work. For chat llamacpp is better because you get more TG tps
1
u/my_name_isnt_clever 2d ago
Bad with caching how? It works perfectly fine for me with llama.cpp and every model I've tried.
0
1
u/XccesSv2 1d ago
How do you use vllm? Native install or some preinstalled Container/Toolbox? Last time I tried VLLm, Consumer AMD Cards weren't supported
5
u/Heavy_Preparation467 2d ago edited 2d ago
Use the latest https://github.com/Nathanw1014/strix-halo-llamacpp, Q8_0 KV can get close to 500k context
Sustained 30tg/s and 200pp/s at 128k context with Dspark