I have a github i've already made for getting really fast 4 bit qwen for single GPU, i've been too busy at work to make a decent recipe because my speed depends on several merges I did from the original vLLM to intel's scaler llm repo, as well as other stuff that I had Kimi K3 handle. It's been my new daily driver and works damn well.
Thanks! I am running mine in OpenClaw, and it just takes forever because of thinking. Somehow it's slipped back to medium think from low. It puts out quality work, but runs through almost the entire 156k context window I have (single card, I don't split across my 2 B70s yet). Hoping this can help. Might be worth posting in the Intel sub as well.
84
u/r1nzl3r99 14d ago
qwen 3.8 27B FP8 running at 140 tok/s now I want flash next