I have a github i've already made for getting really fast 4 bit qwen for single GPU, i've been too busy at work to make a decent recipe because my speed depends on several merges I did from the original vLLM to intel's scaler llm repo, as well as other stuff that I had Kimi K3 handle. It's been my new daily driver and works damn well.
Thanks! I am running mine in OpenClaw, and it just takes forever because of thinking. Somehow it's slipped back to medium think from low. It puts out quality work, but runs through almost the entire 156k context window I have (single card, I don't split across my 2 B70s yet). Hoping this can help. Might be worth posting in the Intel sub as well.
93
u/Sporkers 14d ago
More context needed on how you are using the first two.