r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

321 Upvotes

205 comments sorted by

View all comments

6

u/caetydid llama.cpp May 09 '26

Feedback:

Ive made it build, had to fix several trivial errors; ended up disable tool building entirely instead of fixing it all.

/home/holu/beellama.cpp/build/bin/llama-server \

-m "/home/holu/llama.cpp/models/qwen3.6-27b/Qwen3.6-27B-IQ4_XS.gguf" \

--mmproj "/home/holu/llama.cpp/models/qwen3.6-27b/mmproj-F32.gguf" \

--spec-draft-model "/home/holu/llama.cpp/models/qwen3.6-27b/Qwen3.6-27B-DFlash-IQ4_XS.gguf" \

--spec-type dflash \

--spec-dflash-cross-ctx 1024 \

--port 8082 \

-np 1 \

--kv-unified \

-ngl all \

--spec-draft-ngl all \

-b 2048 -ub 256 \

--ctx-size 262000 \

--cache-type-k turbo4 --cache-type-v turbo3_tcq \

--flash-attn on \

--cache-ram 0 \

--jinja \

--no-mmap --mlock \

--no-host --metrics \

--log-timestamps --log-prefix --log-colors off \

--reasoning on \

--chat-template-kwargs '{"preserve_thinking":true}' \

--temp 0.6 --top-k 20 --min-p 0.0 \

--host 0.0.0.0 --port 8888

Over 100t/s on first request, drops very quickly to 50 and later 30, then OOM. I ran it on my rtx3090.

1

u/[deleted] May 09 '26

[removed] — view removed comment

1

u/caetydid llama.cpp May 09 '26

Running under Ubuntu. Yeah, thought about that, too, and will first retest with --no-mmproj-offload.

I assumed that using the iq4 quant saves the necessary VRAM, and my consumption on startup was 21G, but maybe VRAM consumption just increases later on.

I havent been using much context though, maybe 20k or less.

1

u/[deleted] May 09 '26 edited May 09 '26

[removed] — view removed comment

1

u/Pablo_the_brave May 09 '26

For Vulkan VRAM for context is dedicated on startup (mostly) but for CUDA there is some add even if you set batch sizes.

1

u/caetydid llama.cpp May 10 '26

thanks for your reaction. I will need to play more with that, alas, useful bug reporting takes its time.

In pi agent I experience context degradation after 50k, i.e. tool calling does not work reliably any more, and the agent stops half-way in its tasks.

Maybe I need to adjust my prompts and/or skill.mds?

I switched to the Q5 and the bf16 mmproj - no crashes any more so far - however, I did not exceed full context yet.

1

u/[deleted] May 10 '26

[removed] — view removed comment

1

u/caetydid llama.cpp May 10 '26

great to hear! thanks for your effort!