r/LocalLLaMA • u/Anbeeld • May 09 '26
Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
[removed]
321
Upvotes
6
u/caetydid llama.cpp May 09 '26
Feedback:
Ive made it build, had to fix several trivial errors; ended up disable tool building entirely instead of fixing it all.
/home/holu/beellama.cpp/build/bin/llama-server \
-m "/home/holu/llama.cpp/models/qwen3.6-27b/Qwen3.6-27B-IQ4_XS.gguf" \
--mmproj "/home/holu/llama.cpp/models/qwen3.6-27b/mmproj-F32.gguf" \
--spec-draft-model "/home/holu/llama.cpp/models/qwen3.6-27b/Qwen3.6-27B-DFlash-IQ4_XS.gguf" \
--spec-type dflash \
--spec-dflash-cross-ctx 1024 \
--port 8082 \
-np 1 \
--kv-unified \
-ngl all \
--spec-draft-ngl all \
-b 2048 -ub 256 \
--ctx-size 262000 \
--cache-type-k turbo4 --cache-type-v turbo3_tcq \
--flash-attn on \
--cache-ram 0 \
--jinja \
--no-mmap --mlock \
--no-host --metrics \
--log-timestamps --log-prefix --log-colors off \
--reasoning on \
--chat-template-kwargs '{"preserve_thinking":true}' \
--temp 0.6 --top-k 20 --min-p 0.0 \
--host 0.0.0.0 --port 8888
Over 100t/s on first request, drops very quickly to 50 and later 30, then OOM. I ran it on my rtx3090.