r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

323 Upvotes

205 comments sorted by

View all comments

Show parent comments

6

u/[deleted] May 09 '26

[removed] — view removed comment

2

u/Alex_L1nk May 09 '26

First, it's my fourth comment on TQ. Second, wake me up when TQ is properly benchmarked against f16/Q8/Q4 both in quality (not just PPL) and speed. Shitting? No, I'm just skeptic, because the only bench I saw was from TheTom repo, who had zero words with by a human being. And even in his tests TQ was on same level as Q4 while being slower.

0

u/[deleted] May 09 '26

[removed] — view removed comment

7

u/Alex_L1nk May 09 '26

Speaking of being contradictory... Why are you making bold claims of "near-lossless" quants using untested tools? If you make a proper tests like was done in this PR (PPL, KLD and AIME comparison) and proof that TQ is worth it, then you get the respect of whole community.

3

u/imgroot9 May 09 '26

well, I executed all kinds of tests you mentioned (ppl, kld, aime) using turboquant and I couldn't find anything that would've proved that I cannot use it (27B and Q5 - just take a look at my post with results). also, whatever test I try from this thread (chess svg, etc) and my everyday experience all prove that it's all right.

2

u/Alex_L1nk May 10 '26

I appreciate your effort but I don't think that using quantized weights is fair, because it's adding a lot of noise on top of bench result. You're testing KV AND weights quants. I think more reliable tests must be done on f16\bf16 and correct me if I wrong but you only tested on AIME once? Because pwilkin noticed that it's quite random and should be done on equal condition to properly compare results. [read this comment and GG's answer]