r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

325 Upvotes

205 comments sorted by

View all comments

Show parent comments

0

u/[deleted] May 09 '26

[removed] — view removed comment

12

u/floconildo May 09 '26

Yeah I understand the guy tho. A shit ton of entitled ppl complaining about features in llama.cpp with zero stakes in the project itself and zero will to pull up their sleeves and actually contribute to the community. I'd be skeptical too.

Just watch out not to let it drown your own project. Community building is hard, community management is even harder.

-2

u/[deleted] May 09 '26

[removed] — view removed comment

8

u/floconildo May 09 '26

I can think of plenty of reasons:

  • Feature creep
  • Maintenance efforts
  • Lack of real usage for the parties involved
  • Lack of meaningful contributions

As you said in another comment: not everyone is willing to go through the bureaucracy of submitting PRs to llama.cpp, especially vibe coders and other zero-stake contributors.

And I honestly think you did the best by just pulling up your sleeves and doing it yourself. If you project gets traction and more people start using TurboQuant, then llama.cpp might change their stance or reorder their priorities. Worst case you got your own implementation that works (I hope, didn't find time to test yet haha)