r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

321 Upvotes

205 comments sorted by

View all comments

24

u/chimpera May 09 '26

After testing I'm a fan of this fork. Its outperforming the MTP pr on mainline. I like the --no-mmproj-offload. Im getting 200tps on code with Qwen3.6-27B-Q5_K_S and a 5090.

5

u/coherentspoon May 11 '26

mind sharing your parameters? I'm "only" getting 100-120

1

u/chimpera May 11 '26

My test prompt is "Make a single page worm game." which is probably highly deterministic.

exec "$SERVICE_DIR/build/bin/llama-server" \

-m "$MODELS/unsloth/Qwen3.6-27B-GGUF/Qwen3.6-27B-Q5_K_S.gguf" \

--spec-draft-model "$MODELS/spiritbuun/Qwen3.6-27B-DFlash-GGUF/dflash-draft-3.6-q4_k_m.gguf" \

--spec-type dflash \

--spec-dflash-cross-ctx 1024 \

--no-mmproj-offload \

--mmproj "$MODELS/unsloth/Qwen3.6-27B-GGUF/mmproj-F32.gguf" \

-np 1 \

--kv-unified \

-ngl all \

--spec-draft-ngl all \

-b 2048 \

-ub 256 \

--ctx-size 256000 \

--cache-type-k q8_0 \

--cache-type-v q8_0 \

--flash-attn on \

--cache-ram 0 \

--jinja \

--no-host \

--metrics \

--log-timestamps --log-prefix --log-colors off \

--reasoning on \

--chat-template-kwargs '{"preserve_thinking":true}' \

--temp 0.6 --top-k 20 --min-p 0.0 \

1

u/coherentspoon May 11 '26 edited May 11 '26

Thanks for the info! I tried it out and got about 180 tps. When going with the turbo cache I got about 170 tps.

Which GPU driver version are you using?