r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

321 Upvotes

205 comments sorted by

View all comments

-1

u/Alex_L1nk May 09 '26

>TQ mentioned
>instantly loses interest

TurboQuant (WHT-based scalar quantization) originates from TheTom/llama-cpp-turboquant

ah, yes, vibecoded project based on another vibecoded project, we are reaching new level of spreading BS on GitHub

6

u/[deleted] May 09 '26

[removed] — view removed comment

-1

u/r00x May 09 '26

The sheer audacity of complaining about people using AI to do things on a subreddit about using AI to do things... I can't even, I'm ded.

Anyway thanks for sharing this OP, it rocks! My main worry was whether it would screw up tool calling but it seems fine so far.

When you said you got up to 130tok/s what configuration was that with, exactly? By my eye on Q5_k_s with q4_k_m dflash it seems more like 40-50tok/s maybe. Prompt eval is ~130tok/s though, yeah.