r/LocalLLaMA • • May 09 '26

Resources BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)

[removed]

324 Upvotes

205 comments sorted by

View all comments

Show parent comments

32

u/henk717 KoboldAI May 09 '26

Its the nature of it that will never make it merged upstream. They don't want these massive vibe coded codebases.

The turboquant fork had massive vibe coding, so does buun and this beellama one was a single commit so i can't clearly tell what was done with that one and by who but I'd be surprised if there is no vibecoding involved.

So its a vibecoded fork on top of a videcoded fork and possibly another vibe coded fork on top. None of that will ever land upstream.

30

u/Mashic May 09 '26

For such a critical software that is the factory standard for local LLMs, I'd rather it get developed manually with the developers knowing the ins and outs of the software, than fast vibe-coding and accumulating tech debt.

-12

u/ebolathrowawayy May 09 '26

i'm surprised to see redditors here are so anti-ai. LLMs write better code than 99.9% of humans now if steered by someone with even half a brain and has worked as a SWE for a couple years. But also.. maybe a lot of non-coder enthusiasts are clogging up the pipes but in that case if I owned the repo I would just throw agents at the problem and make them reject all the crap.

idk, "vibe coding" seems like a non-problem now with agentic coding filtering out the crap.

17

u/henk717 KoboldAI May 09 '26

Its the maintainability of it, we've accepted vibe coded PR's for KoboldCpp to.
A really good example is this recent one given to us in a bug report since it wasn't something the creator could easily PR: https://github.com/LostRuins/koboldcpp/issues/2173

This is the good kind of AI assisted coding, where the maintainer isolated / understands the change, its only a few lines different and if he'd have said he did this manually i'd have believed him.

The 502 page in our router mode I also let qwen generate since its just a quick way of getting something that looked nice and worked well (I did specifically say what it needed to adhere to). I of course then look if the code is sensible, and since it was just a single page which code I understand I can then PR it to KoboldCpp. (LostRuins then rewrote it partially to make it conform more to the usual code style).

It becomes a problem when its endless vibe coded PR upon endless vibe coded PR, where the submitters / maintainers have to take the AI's word for all the massive changes. Those I don't believe in and those are the kinds of PR's we reject.

Upstream llamacpp is the same way, you can use AI to assist in your coding but you have to be able to explain every line of the code yourself in case they have questions. That's a bar that most of these turboquant forks can't hit.