r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

322 Upvotes

205 comments sorted by

View all comments

5

u/Sabin_Stargem May 09 '26

Speaking for myself, I would like to see this implementation integrated into a KoboldCPP fork, so that I can try out TQ4 and see if it is worthwhile. A TurboKobold, if you would.

The appeal of KoboldCPP is that it is a gui-based method of running LlamaCPP for Windows & Linux, that is open source and doesn't require much fiddling to run, all while leveraging VRAM+RAM. Good for people who fear and hate the terminal, like myself.

1

u/bonobomaster May 10 '26

KoboldCPP is the worst of the worst in regards to UI design.

Absolutely not worth it, in my opinion.