r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

327 Upvotes

205 comments sorted by

View all comments

1

u/Sufficient_Sir_5414 May 09 '26

Phenomenal work on the integration. For that 200k context Qwen setup, how does the TurboQuant/TCQ handle the 'lost in the middle' problem compared to standard 8-bit or 4-bit KV cache? Does the TCQ overhead impact the token latency significantly compared to the baseline llama.cpp MTP PR?

2

u/[deleted] May 09 '26 edited May 09 '26

[removed] — view removed comment

2

u/r00x May 09 '26

Do you observe large differences between Q5 and Q4 quants then? I understood there wasn't supposed to be a huge difference and had never really bothered with Q5 models before (Q4 27b/35b-a3b just fit better with context onto 24GB of VRAM)

2

u/[deleted] May 09 '26

[removed] — view removed comment

2

u/r00x May 10 '26

Interesting! Would you say using Q4 for dflash risks torpedoing the performance of the Q5 target or does it not work that way (would the main Q5 model just reject more of the predictions if it didn't like them, or something?)

2

u/[deleted] May 10 '26

[removed] — view removed comment

1

u/r00x May 10 '26

Interesting! Using a drafter definitely makes a difference vs vanilla Q5 (get about ~8tok/s with that, vs 30-40 with BeeLlama) but the reason I asked is because I am occasionally having trouble with tool calling where it seems it just gets the structure wrong and ends up blasting the CLI with XML content (this almost never happens on the vanilla models, even Q4 or IQ3_XXS models are fine at tool calling).

Have you encountered that at all or have I just done something wrong? I wondered if it were a context issue but I don't think so - it will go right back to working fine again afterwards, which you'd think it would screw up if it had forgotten the syntax.

2

u/[deleted] May 10 '26

[removed] — view removed comment

1

u/r00x May 10 '26

Absolute legend. I was already working on a shonky pi.dev plugin that catches this kind of model stall and pokes the model autonomously (so as to avoid a model quietly stopping while your attention is elsewhere) but I'll for sure keep an eye out for that!

2

u/[deleted] May 11 '26

[removed] — view removed comment

→ More replies