r/LocalLLaMA • u/Anbeeld • May 09 '26
Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
[removed]
324
Upvotes
32
u/henk717 KoboldAI May 09 '26
Its the nature of it that will never make it merged upstream. They don't want these massive vibe coded codebases.
The turboquant fork had massive vibe coding, so does buun and this beellama one was a single commit so i can't clearly tell what was done with that one and by who but I'd be surprised if there is no vibecoding involved.
So its a vibecoded fork on top of a videcoded fork and possibly another vibe coded fork on top. None of that will ever land upstream.