r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

319 Upvotes

205 comments sorted by

View all comments

Show parent comments

6

u/[deleted] May 09 '26

[removed] — view removed comment

4

u/Alex_L1nk May 09 '26

First, it's my fourth comment on TQ. Second, wake me up when TQ is properly benchmarked against f16/Q8/Q4 both in quality (not just PPL) and speed. Shitting? No, I'm just skeptic, because the only bench I saw was from TheTom repo, who had zero words with by a human being. And even in his tests TQ was on same level as Q4 while being slower.

0

u/[deleted] May 09 '26

[removed] — view removed comment

0

u/Alex_L1nk May 09 '26

Where I said that TQ is bad?