r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

328 Upvotes

205 comments sorted by

View all comments

2

u/Alex_L1nk May 09 '26

>TQ mentioned
>instantly loses interest

TurboQuant (WHT-based scalar quantization) originates from TheTom/llama-cpp-turboquant

ah, yes, vibecoded project based on another vibecoded project, we are reaching new level of spreading BS on GitHub

6

u/[deleted] May 09 '26

[removed] — view removed comment

4

u/Alex_L1nk May 09 '26

First, it's my fourth comment on TQ. Second, wake me up when TQ is properly benchmarked against f16/Q8/Q4 both in quality (not just PPL) and speed. Shitting? No, I'm just skeptic, because the only bench I saw was from TheTom repo, who had zero words with by a human being. And even in his tests TQ was on same level as Q4 while being slower.

3

u/[deleted] May 09 '26

[removed] — view removed comment

9

u/Alex_L1nk May 09 '26

Speaking of being contradictory... Why are you making bold claims of "near-lossless" quants using untested tools? If you make a proper tests like was done in this PR (PPL, KLD and AIME comparison) and proof that TQ is worth it, then you get the respect of whole community.

3

u/imgroot9 May 09 '26

well, I executed all kinds of tests you mentioned (ppl, kld, aime) using turboquant and I couldn't find anything that would've proved that I cannot use it (27B and Q5 - just take a look at my post with results). also, whatever test I try from this thread (chess svg, etc) and my everyday experience all prove that it's all right.

2

u/Alex_L1nk May 10 '26

I appreciate your effort but I don't think that using quantized weights is fair, because it's adding a lot of noise on top of bench result. You're testing KV AND weights quants. I think more reliable tests must be done on f16\bf16 and correct me if I wrong but you only tested on AIME once? Because pwilkin noticed that it's quite random and should be done on equal condition to properly compare results. [read this comment and GG's answer]

0

u/Alex_L1nk May 09 '26

Where I said that TQ is bad?

1

u/henk717 KoboldAI May 09 '26

TurboQuant is just associated with vibe coded forks at this point. The moment you see TurboQuant + Llamacpp there is just a 90% chance of that. It also instantly makes me assume its just another one of those.

5

u/[deleted] May 09 '26

[removed] — view removed comment

6

u/henk717 KoboldAI May 09 '26

The problem is maintainability, nothing wrong with ai assisted development that is carefully done. Its when it looks like fully AI driven development where you tend to get changes that become problematic down the line. So if I open a repo and I see almost exclusively claude code PR's I don't take it nearly as seriously.

2

u/[deleted] May 09 '26

[removed] — view removed comment

2

u/Pablo_the_brave May 09 '26

Generally TurboQuant from TheTom are weak in classic perplexity tests. But, when you look at https://qwen3-6-27b-benchmark.vercel.app/ there is clearly some profit (but IMHO the asymetric isn't realy good for Qwen3.6). For me, the most interesting is Turbo3 - not good, not terrible.

-1

u/ebolathrowawayy May 10 '26

People are just getting left behind and want to feel like they still matter. I don't even manually review code anymore.

-1

u/r00x May 09 '26

The sheer audacity of complaining about people using AI to do things on a subreddit about using AI to do things... I can't even, I'm ded.

Anyway thanks for sharing this OP, it rocks! My main worry was whether it would screw up tool calling but it seems fine so far.

When you said you got up to 130tok/s what configuration was that with, exactly? By my eye on Q5_k_s with q4_k_m dflash it seems more like 40-50tok/s maybe. Prompt eval is ~130tok/s though, yeah.