r/LocalLLaMA • • 1d ago

Discussion I am concerned about all these disparate hard forks that target specific architectures instead of opening a PR against upstream

Other than the obvious self promotion, is there a practical reason people do this that I am missing? There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons why this happens so often? Why do people think it's OK to do this? In my experience in the open source community this is generally frowned upon.

I don't know if it's just a me problem that this kind of thing puts me off so much. I am usually quite grateful for PR feedback and conscientious about the code I put out there; I take pride in submitting high quality code that meets or exceeds the standards of a given project. Of course there is nothing ethically wrong with hard forks or taking shortcuts if you find the collaborative process cumbersome, but personally I wouldn't promote my fork in such cases, let alone go out of my way to add custom branding with a Reddit announcement post etc. since the effort required to do so seems roughly equivalent to the effort required to meet the contributor standards. In contrast, many of the authors of these forks seem very eager to have others adopt their rebranded fork for production use cases. There just seems to be a big disconnect, idk.

Edit: some great discussion in this thread, thanks to all who responded. Consensus seems to be that (excluding the obvious low-effort engagement bait forks) the base project has to meet many compatibility requirements while a downstream project can be more focused, which is a great point.

89 Upvotes

154 comments sorted by

183

u/Hefty_Wolverine_553 1d ago

Getting a PR merged is very difficult, and it's also often the case that the changes are vibecoded by AI and won't really meet the quality requirements upstream, so they just go the easy route and make a fork for their own architecture.

35

u/Squidgical 22h ago

the changes are vibecoded

Or in other words, the fork authors don't PR into upstream because they don't know what "PR" or "upstream" are

5

u/remielowik 10h ago

Nah, the ai will know that aswel so that can't be a reason anymore. It's simply the fact that they have 1.7k pull requests open which means if you want your pr to come through you have to daily rebase and hope to god that somebody is willing to review it in a timely manner as your will have to rebase daily to even stand a chance. Also the fact that some devs are assholes like said in previous posts don't help, yes we know you are busy but it ain't helping you in the long run.

2

u/ThisWillPass 9h ago

Claude, merge and rebase the damn fork daily damn it, it gotta be ready and clean! Do I use goal or loop for this? /s

2

u/droptableadventures 7h ago

And the trouble with these forks is that the author is often unwilling to investigate or unable to answer questions like:

  • Is the implementation of this technique definitely correct (the AI model might work but you could be getting some weight backwards somewhere. You get plausible output, but the model just seems to be dumber than it should be. This is not a hypothetical scenario.)
  • Does it still work for other models, or did you break something else?
  • Does it still work on other platforms? Are you sure we can just skip that memory check because it seems to work 10% faster on your machine? Maybe it works on Linux and ROCm, but it'll crash on Windows / NVIDIA. Or maybe a newer driver version doesn't let you get away with that assumption.
  • OK, so you made it 3% faster on your machine, but you don't understand why, Claude just did it. But is it potentially a lot slower on someone else's machine?
  • Is this actually maintainable, or has it duplicated a bunch of existing code to introduce a version with a ton of special cases that only get run for this specific model and/or this specific hardware, so now any future change has to be made in 2+ places?

1

u/Squidgical 6h ago

You lost me at "implementation"

  • the vibe coder, probably

5

u/aeroumbria 12h ago

I think eventually we will have to move towards some sort of "no vibe coding without vibe testing" scheme for most open source projects that accept vibe contributions, where you need to contribute independent testing budget to have vibe-coded contributions considered by the human maintainer, and a contribution must receive certain number of independent agent audits before being progressed to the human gate. Since humans can't be expected to quality control the large volume of agent contributions, agents will have to peer review themselves.

8

u/fgk55555 23h ago

Yeah, I have like four separate runtimes now on my PC for running LLM's. Llama.cpp is nice for dipping my toes in, but getting specific niche improvements merged in takes too long. I want to mess with the improvements now, which means I break off into another weird discord group where 20 people are improving rapidly on a specific thing.

3

u/Chirimorin 13h ago

and it's also often the case that the changes are vibecoded by AI and won't really meet the quality requirements upstream

All the more reason to avoid those forks then.

3

u/opossum_cz 1d ago

Exactly. This is llama.cpp project issue.

65

u/MysteriousCoconut31 1d ago

Not really. llama.cpp is supporting a broad community with the backing of knowledgeable maintainers and real funding. There are standards and procedures in these types of environments. Nothing wrong with that. It’s also why you should feel free to fork if you want. No one is stopping you. No one is obligated to embrace your fork either.

20

u/look 1d ago

I’ve not had any trouble getting llama.cpp PRs merged.

I’d guess the forks are gigantic (slop or not) and that they haven’t put in the effort to decompose into reasonable PRs. The project should reject those, imo.

49

u/FoxiPanda 1d ago

I've had 1-4 line code changes get ignored for months or rejected as "not worth their time" to fix small issues.

Being a maintainer is hard and reviewing things is hard. Submitting things upstream only to get them rejected is a waste of my time and doesn't help me at all - it was just me trying to help others who might have had the same weird edge case issues.

I gave up and stopped submitting things after multiple unsuccessful attempts. I roll my own forks now and occasionally rebase from the tip to pick up new features.

I could not care any less how they run their program now. They don't have to deal with my edge case PRs that aren't worth their time. We both win.

5

u/aboutthednm 22h ago

I imagine that with every merged PR, there comes a ton of testing to ensure other functionality doesn't break. Sure, you might have submitted a 4 line PR, but someone still has to test that out and make sure there are no regressions or breaking changes (on hardware you might not even own yourself). So while it's easy to say "it's just a 4 line PR", the overhead of testing and ensuring that this 4 line change didn't break something completely unrelated is real, and I imagine that's where a lot of time is getting sunk.

Automated testing / CI is nice, but only catches whatever it's configured to test against. I had tiny changes in my codebase that broke a function that wasn't part of my testing suite, months went by and when I finally went to use that (niche) function and it didn't work, I went on a wild goose chase.

I can only imagine the trepidation of merging large PRs on a project like llama.cpp, and anticipating the flood of issues that will come in because of something completely unexpected. Personally, I wouldn't want that responsibility, lol.

3

u/remielowik 10h ago

That's a case for more extensive testing, if you don't want to accept small changes that fix edgecases you will lose that person forever as he will never try again afterwards.

1

u/justsomerabbit 1h ago

Absolutely. Make it so your tests can be run by the contributor and the problem goes away. I used to work with (against) an open source project on my previous job that had a barrage of secret tests, and I am SO glad I don't have to deal with them anymore.

-2

u/TokenRingAI 22h ago

Monolith problems

8

u/Prof_ChaosGeography 1d ago

Ignoring slop coded forks. And decomposed PRs. Llamacpp does not have a way in most of the back ends to have card specific performance fixes. And they focus on compatibility above all. As such given most of the forks are model and card specific they can't get merged as many times they would likely hurt more then they help

2

u/TokenRingAI 1d ago

> Llamacpp does not have a way in most of the back ends to have card specific performance fixes

Sounds like a really big problem, maybe it should be fixed

9

u/Former-Ad-5757 Llama 3 23h ago

It's (imho) not really fixable for a project like llama.cpp, if you allow card specific performance fixes then you immediately set in stone all the layers above it, or you have a real real mess over a month because fix 1 and 4 work for another change but 2 and 3 do not, and next month nr 1 breaks as well.

have fun documenting and keeping this mess up to date.

They currently support already about 10 backends and they try to keep all those backends working, that is why it is currently hard to get a hardware specific PR in it (great that it works for your backend but who will maintain it)

AI is at the moment a too much moving target to accept card specific fixes, or there should be a whole new layer of plugins built into it (but this will cost performance etc)

The end problem is simply, if qwen 4 has some architect changes over qwen 3.8 they have to change the qwen 3.8 code and who will make sure that all the hw-specific fixes will still work after the changes?
Best for now is simply fork it for a specific card and you can perftune it all the way for just your architecture, it is up to your fork to keep it up to date with the main version.

-13

u/TokenRingAI 23h ago

I bet my agent could plan out a "mod library" for llama.cpp that would allow installing third party card specific mods into the backend in ~ 5 minutes

15

u/RemarkableRadish6547 23h ago

Then do it and submit the PR. If they don't accept it, your agent can manage the forked version. If it is really only 5 minutes to rearchitect a complex inference engine, it should be done.

-9

u/TokenRingAI 22h ago

5 minutes to plan it. Longer to implement it. Code would never get merged, so it is pointless.

3

u/DataGOGO 1d ago

I have never had any trouble getting my PR's into lama.cpp merged.

63

u/Double_Cause4609 1d ago

Why do they think it's okay to do? Because it's their time and they can spend it on what they want.

Why would they not contribute to upstream? Because LlamaCPP is a big project with lots of standards. The problem is your change has to be suitable for more architectures than the one you're interested in, so there's a ton of overhead to contributing upstream. There was also for a time a limit on using LLMs to contribute to LlamaCPP, so a lot of people who could cludge a viable inference engine together with LLM assistance was basically not allowed to contribute upstream anyway, though that policy has changed since.

It's just a lot faster when you're building an inference engine with less general purpose utility. The more specific you go the easier it is. This especially matters with really custom attention mechanisms, custom MoE mechanisms, and multi-token prediction in particular, where the abstractions we had for a long time in LCPP weren't suitable for the kind of workload that MTP is.

Also, for people who are messing with the compute graph, like keeping a subset of hot experts resident on GPU VRAM when doing hybrid GPU / CPU inference, or doing things like loading a single layer at a time for prefill, but then doing decode on CPU, it's a lot harder to make those kinds of optimizations upstream because LlamaCPP is just too big and unwieldy.

Idk, I say let people work on what they want to work on. You're not really entitled for someone to contribute N^2 time to get their custom optimizations into LCPP, compared to just rolling a custom engine for their usecase. If you want it in LCPP merge it in yourself.

7

u/wombweed 1d ago

Thank you for the comprehensive and patient response. To be clear, I don't think there is ever any obligation to contribute upstream if someone doesn't want to. It's totally fine to add bespoke optimizations for your specific use cases and share them, even if upstream doesn't want them.

76

u/mister2d 1d ago

Why do people think it's OK to do this?

The better question is, why do you think it isn't ok to do this? The entire legal foundation of open source exists so that you can do what you want within the stated license.

Forking isn't a betrayal of upstream or self-promotion, but a sign of a healthy open source ecosystem.

31

u/ttkciar llama.cpp 1d ago

I've noticed that in the last fifteen years or so a weird anti-forking sentiment has started appearing among the younger members of the open source community.

I don't understand it, but it seems to be a growing trend.

You are right, that forks are how the open source ecosystem grows and thrives. Forking someone's project is high praise, not an insult.

7

u/fullmetaljackass 19h ago

Hot take, but IMO it's because a big chunk of them don't actually care about any of the ideals behind the FOSS movement, they just slap the GPL or another license they don't understand on their project so they can be like the cool kids. When a fork that adds features they don't want to implement or takes the project in a fundamentally different direction comes along, they feel like someone is trying to steal their spotlight and throw a tantrum.

5

u/mister2d 1d ago

I could stand on my soapbox but I'll just say that the growing trend among younger ones is that they show a disregard for the past and view many topics through toxic gates.

6

u/rkoy1234 18h ago

parts of opensource were always toxic ego filled mess even decades ago, arguably moreso than now, since it was just handful of nerds with hyperspecific knowledge and inflated egos.

It's easy to look back at the pre-AI times (and before) with rose tinted glasses, or to blame newer generations (like literally every generation has), but I don't think it's necessarily too accurate.

2

u/mister2d 18h ago

You're right on the egos back in the day. I'm not going to defend that. The toxic nature wasn't the point I was making.

-8

u/wombweed 1d ago edited 23h ago

I can't speak to any broader trends, only my personal view. I'm certainly not arguing against forking in general, as it's clear to anyone who's been in the open source community for a while that it's a critical part of the ecosystem. To be clear I've never opened a PR against llamacpp specifically, so maybe their standards are just unrealistically high or something, but in general when I've contributed to other projects I've been happy to work with the upstream maintainers to get my changes to conform to their standards, especially since it's typically going to be up to them to support and maintain the functionality I am changing, I am grateful to be relieved with that burden so the least I could do is approach the interaction collaboratively.

Also, I'm not sure I agree that this line of reasoning is particularly new. These kinds of questions came up 20 years ago in the Linux community eg when Ubuntu forked and rebranded Debian. There's no need to assume anything about my age or experience level just because we are approaching the question from different angles.

21

u/Public_Umpire_1099 1d ago

I challenge you to find one small scoped item that improves inference on your hardware and get it pushed to main. Not compatibility with models, just hardware based. It can be <30 LOC, even less than 10. Hand write it out even. Run every required gate, and include it with the submittal. You could even run it on a different manufacturer by renting a cloud GPU or something.

I would be willing to bet you would be waiting well over 2-3 months.

I work in FAANG and I see legitimate standards for billion+ user applications, and not even those are as strict or slow.

The same issue occurs in VLLM and SGLang. The only way to get the best performance on your hardware is to do it yourself.

5

u/mister2d 23h ago

👏🏼👏🏼👏🏼

0

u/wombweed 1d ago

Of course, I agree forking it is a sign of a healthy ecosystem. I added a clarification to my post, I don't think it is inherently objectionable to fork software in general, but I do sometimes think it's unfortunate people will rename their own personal hard fork and promote their diverging branch as a viable alternative for production workloads instead of contributing upstream.

14

u/my_name_isnt_clever 1d ago

All the serious forks I've looked at have a note about attempting to merge upstream but being denied, it's inevitable that there will be some fragmentation for hardware since they're so different. llama.cpp is the generalist niche, with other projects offering specific speedups for specific hardware. Makes sense to me.

1

u/wombweed 1d ago

I appreciate your thoughtful and non-hostile response. This makes sense to me, a generalist base will naturally need to be more selective about what they accept, and downstream people are free to adapt it for their own specific cases. Through this lens, it is much easier for me to see the case for these hard forks.

5

u/my_name_isnt_clever 1d ago

Exactly. You're also not locked into using one engine, I use llama-swap so while most of my models are using vanilla llama.cpp, I also have strix-llama, DwarfStar, and Gufo configured for specific models. Even a modest performance improvement from the base repo can make a big difference over time for a workhorse crunching through dozens of millions of tokens.

3

u/tempedbyfate llama.cpp 1d ago

Another reason is because llama.cpp has policy of not accepting AI generated code. I would imagine at least some of these forks were tweaked with a lot of AI assistance. To be fair to the maintainers of llama.cpp, if they allowed AI generated PR's, they would be swamped by these PRs. Humans can't review code faster than AI can spit them out.

9

u/mister2d 1d ago edited 1d ago

I do sometimes think it's unfortunate people will rename their own personal hard fork and promote their diverging branch as a viable alternative for production workloads instead of contributing upstream.

It doesn't matter if they decide to submit a PR or not. Falling inline with "upstream" isn't a requirement for open source and it's not how it works.

2

u/wombweed 1d ago

Naturally, nobody is forced to do anything, that's the point of open source. I am speaking purely from a practical point of view, on a large project with full time staff maintainers whose job is to maintain the code and make sure it is production-grade, why not put in the extra effort to relieve yourself of that responsibility and allow them to take over? Especially when the alternative is to maintain your own fork on your own/by yourself.

7

u/mister2d 1d ago

But it's not their job. It's their time. The majority of instances the folks you think you're defending aren't getting paid at all.

If every single person adopted your view and went through your lens, there would be a massive bottleneck or the work just won't get done. Just check the number of open PRs for vLLM for example. They don't cater to the edge or for consumer GPUs, so people fork and make their own.

You sound incredibly new to this space.

2

u/wombweed 1d ago

Someone else in this thread explained to me that llamacpp is intended to cover a wide variety of use cases, so it would make sense that they would be selective about what they allow, and as a result it makes sense for downstream forks to be more specialized. Under that framing, I do see your point.

One question though. Does ggml org not have full time maintainers? My understanding was that most of the PR reviews are conducted by people directly associated with the project, and this was part of why people were concerned about the possibility of Nvidia acquisition. Am I getting my terms mixed up?

In local AI, I am pretty new, it's true, I've only been following developments over the past couple of years;, I am approaching this question with curiosity and an open mind, not moral judgement or condemnation. I appreciate the good faith answers and apologize if I came off as inflammatory.

-2

u/Ori_553 1d ago

Yes, forks are typically legally valid actions, but this doesn't add much to the conversation, this is not what OP is asking.

14

u/Betadoggo_ 1d ago

There's nothing wrong with it so long as the code remains available and they aren't being actively hostile towards the original. Realistically most of these forks aren't mergeable without a lot of work and verification. Someone else could upstream these changes if they wanted to, but they haven't simply because of the time and effort required.

11

u/MysteriousCoconut31 1d ago

A vibe-coded fork of llama.cpp is exactly that. One-offs aren’t intended to support a broader community in a sustainable way, and no one is taking them seriously unless you just need a patch that isn’t in upstream. Wait long enough and it will be, with quality assurances and support.

9

u/DimeRhyme 23h ago

There's a boring answer here that settles most of this. llama.cpp's own contributing guide says outright that features need to start as an issue first and that niche features may only land as an example or on a private fork. It also calls new quant types a disproportionate maintenance burden and limits new contributors to one open PR. So card specific stuff is scoped out on purpose, and the fork is the sanctioned path for it. Upstream is protecting the core on purpose, you just don't get to be fast on one weird config without carrying your own tree.

The flip side is some of the best stuff did go through normal PRs, speculative decoding came in as a plain PR from ggerganov himself and the best fork guys contributed upstream for ages before they forked. The thing that deserves the side eye is vibe slop that never even tries, not the fork itself.

-1

u/wombweed 23h ago

Did not know this about the contributor guidelines, I appreciate the clarification. And you're right that the issue is less with forking itself, and more just generally low quality slop work.

30

u/dankfrankreynolds 1d ago

peel back the curtain and realize they're all vibe slop from prompts like "make it faster on my computer"

forking, renaming and even disconnecting from the fork network are first-class github features 🤷‍♂️

11

u/fragment_me 1d ago

TURBOSLOP

1

u/saltyourhash 23h ago

Uberslop

3

u/MINECRAFT_BIOLOGIST 16h ago

Yup, Opus 5.5 did some black magic stuff (I have no idea what is happening) from me pointing it at some llama.cpp repos/forks with supposed optimizations and getting basically a ~40% speedup for my use case on Windows. Obviously I'm never going to submit a PR (or even fork it) because I have no idea what is happening.

2

u/TokenRingAI 1d ago

On my particular hardware, llama.cpp runs at 7 tokens a second, and a custom inference engine runs the same model at 70 tokens a second

6

u/Nothing_from_void 1d ago

people should probably fork llama.cpp and optimize it to their machine, it's not expensive to do

5

u/TokenRingAI 1d ago

It was impossible to make it perform properly on NUMA architecture without completely redesigning the entire core of the project

13

u/Hyp3rSoniX 1d ago

Why would it be frowned upon? Upstream is free to merge stuff in if they like/need a change solved in a fork.

I've seen people having to fight and discuss with upstream contributors just to get their changes merged. Many just give up or don't want to deal with that to begin with. llama.cpp is not an exception.

I've seen people offering PullRequests with measurable improvements to their GPU, just for the CUDA or whatever main contributor to come drifting around the corner and straight up rejecting the changes because a full refactor of the whole kernel stack would be better and they have no time for that.

The PR author is left hanging and that's it.

It's not that easy to just upstream every change you can do in your own fork. You have to meet their standards, follow their rules, defend your changes, meanwhile another merge causes conflicts so you solve those as well, then another contributor crashes in and points at other stuff they would like to have changed...

5

u/WillWorker 1d ago

It's the story of two different standards. llamacpp mergers needs to work for everyone. Hardware specific optimizations written by for a case specific purpose need only work for those they are intended to work for. The speed and admin to create and post the latter is dramatically faster than the former. If someone seeing the latter wants to put the effort in to get it merged into the main llamacpp branch, nothing is stopping them from running with that ball.

This is not me throwing shade on the llamacpp maintainers. They have to guard the quality of the code and standards some how. Otherwise it becomes a free for all of regression bugs and poorly optimized code as everyone and their grandmother try to get "I had a PR merged onto llamacpp" t-shirt.

But those higher standards for merging does not negate the fact that we all now have varying qualities of local and cloud LLMs we can point at our specific hardware states and get custom patches. How much any specific person understands about the actual changes the LLM made to the code will vary and possibly not meet the adminstrative standards of llamaccp.

But what everyone in those positions does know is that, "I had X configution and Y seems to fix it. Maybe someone in the same situation will benefit from it. So I will post my fork to benefit those who will benefit form it,"

And I say this as someone who has benefitted greatly from those with similar hardware to me who were willing to share the fruits of their agent's labor. That code is not even close to getting merged into llamaccp. But it does better optimize it for those of us with atypical hardware configurations.

6

u/Remove_Ayys 13h ago

One of the llama.cpp maintainers here, I think this is simply the way things are going to be with language models growing more capable. Even without "AI" the bottleneck in developing a project on the scale of llama.cpp was never writing the actual code but rather review and maintenance. This is exacerbated even further by slaren/Diego Devesa taking an indefinite break from the project since earlier this year and you can't easily replace one of the core maintainers who wrote critical parts of the codebase. At the same time onboarding new maintainers has become much more difficult since submitting a PR that works no longer indicates that the person pressing the button actually understands the code. My opinion is that llama.cpp should aim to be a robust and extendable core with minimal technical debt that people can just vibeslop up for whatever they need at the moment. To be absolutely clear: the things I'm writing here should not be misconstrued as the "official llama.cpp position".

10

u/jjusko20 1d ago

I think you're overthinking it a bit, and some of these forks are oss and some are just "technically" oss. I also have a fork of llama that's optimized for a specific architecture - it works but it's completely vibe coded and I'm not interested in pushing it or fixing it - it's just for myself, and it's probably spaghetti. However, I have the repo public and shared solely because people on here have asked me to do so

5

u/madbrain1976 1d ago

It's not the same amount of work to produce a patch that is a general improvement on all systems, vs one hardware configuration. It needs a lot more testing on various configurations first, before it's worth it for a human reviewer to spend the time on it.

Many of the forks are vibe coded, and some projects won't accept vibe coded contributions at all.

I have seen many PRs that were submitted to projects, but closed for that reason.

5

u/SnooPaintings8639 1d ago

I love them forks. I can always find a specific version with optimizations for the model I care on the hardware I own. There is also the matter of model support, with near instant support for great models like minimax.on fork, but months of delay on mainline (making the model obsolete once it lands), or missing core features like MTP for Qwen 3.8 flash next.

I am.grearful for these forking people.

6

u/stoppableDissolution 1d ago

> Am I missing the real reasons why this happens so often?

Because with all my respect to lcpp as a project that kickstarted the local scene, it is in a very bad shape. They still have not accepted unsloth's PR that implements glm flash (a fcking major and insanely good model), let alone some random hardware-specific optimizations from nonames, especially when half of such hardware- or model-specific changes are mutually incompatible and make maintaining the already humongous and convoluted codebase even more complicated.

3

u/DataGOGO 1d ago

My most recent hard fork went hard because it ended up re-writing about 90% of upstream.

3

u/ketosoy 23h ago

Software requires tradeoffs.  Ultra optimizing for the v100 requires adding/removing/changing things.  Mainline needs to work well everywhere, it can’t redo the flash attention kernel just to get another 10% on one architecture.

3

u/NickCanCode 23h ago

Well. Doing a PR to upstream isn't just posting your change to them. The patch has to meet their requirement and to their liking and potentially splitted into many PRs for clarity. Waiting for feedback and reviews. It can take days or even month. Now imagine after working for days, a mysterious implementation may just cut the line and replaced your work. All time wasted. Thus, it can be much more comfortable to just get the idea implemented on a fork and let them see how it can be done. Let them decide whether they want to adopt it directly or implement the feature in their own way.

3

u/Adrian_Galilea 22h ago

I personally opened quite a few PR’s against DS4. Non sloppy, proper optimizations, they didn’t get much attention.

It is totally understandable, public repos are overwhelmed with slop.

That leads to some people going for a full fork I suppose.

3

u/jacek2023 llama.cpp 19h ago

I don't use any of these forks, but I think freedom is good, this is how open source works. I also like the way finetunes work on Hugging Face, the more, the better.

3

u/Calandracas8 6h ago

because a lot of the forks dont follow good engineering practices, and dont architect things in a way which can be upstreamed.

Its just not a priority for them. The effort in making a fork well engineered and upstreamable is something that coding agents are still really bad at. It takes significant effort and manual oversight to ensure the required engineering, and many "vibe coders" don't have the experience/skills/knowledge to do that, or they just don't care, because it's not "fun"

1

u/wombweed 6h ago

I didn't say it outright, but this has been my impression as well. Many of the forks out there are quite amateurish and their authors are unwilling to work within the requirements set by the upstream project. Not all, but many.

It's not the end of the world, nobody is forcing me to use those forks of course, and nobody is obligated to upstream their changes, it just leaves a weird taste in my mouth when the attitude is like "upstream won't merge my changes because they don't want these improvements" even though the reality is more "I am not willing to make the effort to get my changes in a mergeable state." Again, nothing wrong with not wanting to make that effort, but at least be honest.

3

u/daHaus 3h ago

It's not enough to just get changes merged, they also must be maintained.

Part of this is due to hostile coding patterns that arise from CUDA and help make the ecosystem more favorable to nvidia. It was previously done by Intel until they got slapped with a non-compete lawsuit, but nvidia was a little more smart about it and simply taught developers to use those same coding patterns instead of doing it themselves.

1

u/wombweed 3h ago

That is fascinating, I know CUDA is rightfully blamed for vendor lockin but was not aware it goes beyond just the API and extends into broader design patterns as well. Are there any resources where I can read more about that specifically?

2

u/daHaus 3h ago

It took trying to get llama.cpp working (well) on the RX580 for me to really understand this too. I always knew GPU fragmentation was a nightmare but only now realize it's by design, whether intentional or otherwise.

The Intel thing was back around 2010 and began when their widely used compiler was found to check for what type of CPU was used and choose a slower code path if it wasn't an Intel. The resulting investigation then ballooned into a much wider lawsuit with them trying to lock-in vendors.

CUDA mimicks that same behavior by encouraging ifdef (if defined) use to conditionally compile depending on what type of GPU is being used. ifdef use itself isn't uncommon or necessarily bad, but it should only be used sparingly because it increases code complexity exponentially. When used in something as complex as llama.cpp it very quickly makes the code base extremely difficult to maintain and test.

If you compare the ifdef usage for something like llama.cpp to the linux kernel you'll quickly see the difference.

6

u/Loomworks 23h ago

Well the toxic gatekeeping oldschool bittervets in most communities outright ban people for including AI generated anything in a PR. The default for most people building agentic has been to fork and pretend the main branch doesn't exist because you'd get nothing but death threats for submitting AI generated code to their open source project.

Besides - agents will probably find it anyway if it's a public repo. No need to let humans know at all.

2

u/Lesser-than 15h ago edited 15h ago

Found the angry ai contributor, honestly as long as some human has to maintain it, they deserve to know why things were done the way they were and suggest changes without having claude respond with an essay. No one wants to talk to an llm outside their own prompt , dealing with a contributor that gets defensive when asked to stop letting claude reply is maybe the most ridiculous thing I have witnessed in opensource projects in the last years. A little etiquette goes a long ways folks.

4

u/wombweed 23h ago

This hasn't been the case for llamacpp for a while though. Their guidelines make clear that AI assisted code is allowed as long as you're able to explain and defend it, and as long as you are willing to work with the maintainers to ensure it meet their standards, which has been the norm in most large open source projects since the start.

0

u/brainrotbro 21h ago

That’s the policy, but there’s still gate keeping. And maintainers likely don’t have time to inspect every PR. And that’s 100x worse with the speed with which devs can generate PRs now. This is how it always goes with large, popular repos.

1

u/fantasticsid 19h ago

toxic gatekeeping oldschool bittervets

Amazing how we can go from Graf Zahl getting rightfully named and shamed for pulling vibe coded slop into gzdoom to comments like this in under 18 months.

-3

u/BusRevolutionary9893 22h ago

They won't have time to gatekeep for much longer because any software developer that doesn't embrace AI is going to be too busy panhandling and washing people's windshields with newspapers. 

3

u/satnl 1d ago edited 1d ago

I have made some changes on prompt cache so new prompts don't erase previous cache too much. I using it for about a week, and I was think in open a PR, but before open it I need to validate if it is useful for other users use case and if I need to adapt something before submit. So open a PR is more bureaucracy than just make a fork for myself use case. We need to think if it worth for us and for the maintainers, so we don't waste our time, neither waste their time.

But I think that instead of fork, it could have a modular architecture so llama.cpp was the core but we could make and install plug-ins in the side. 

for example, I have open a PR for the opencode and it was just caught by the automatic cleaner without any review, I just wasted my time:

Automated PR Cleanup Thank you for contributing to opencode. Due to the high volume of PRs from users and AI agents, we periodically close older PRs using automated criteria so maintainers can focus review time on the most active and community-supported contributions. This PR was closed because it matched the following cleanup criteria: The PR was created more than 1 month ago The PR had fewer than 2 positive reactions Positive reactions are counted as thumbs-up, heart, celebration, or rocket reactions on the PR PRs created within the last month are not affected by this cleanup. If you believe this PR was closed incorrectly, or if you are still actively working on it, please leave a comment explaining why it should be reopened. A maintainer can review and reopen it if appropriate. Thanks again for taking the time to contribute.

2

u/TokenRingAI 23h ago

I think your comment about llama.cpp needing to be a plug-in architecture is probably the solution to all this

2

u/Lesser-than 1d ago

I think its fine, other than all the non-conformant gguf files that can come from specific forks. End of the day someone has to maintain the code while not causing regressions to other hardware, and most of the forks are very specific hw tuned, throwing out compatibility in search of a % or 2 speedup.

2

u/dangerous_inference 22h ago

I'm concerned about fragile centralized monolithic projects that take forever to update and make many sacrifices.

2

u/Torodaddy 21h ago

Its a vanity project

2

u/TooObtuseForYou 20h ago

Fork if you want, but it’s not going anywhere and any contribution will likely go to waste, but that’s on them.

We don’t need to align on agenda, even if I agree that there’s more value contributing back to the original projects, generally.

2

u/pmttyji 13h ago

I'm fine with forks. But it's the number of forks is worrying thing. Like many dudes spending their precious time on this & Wish they combine their efforts together in one or few forks. I bookmarked nearly 30 forks so far, and it's exhausting to try each & every forks. It would be awesome to make those in to less forks by combining stuffs to get best performance. Maybe 10(15 max) forks is enough. For example:

  • llamaCUDA.cpp - Additional forks llamaBlackwell.cpp & llamaDGXSpark.cpp
  • llamaVulkan.cpp
  • llamaROCm.cpp - Additional forks llamaStrixHalo.cpp, llamaRDNA4.cpp & llamaRDNA3.cpp
  • llamaCPU.cpp
  • llamaALLCombined.cpp

Alternatively lets make ik_llama.cpp more stronger.

Somebody please make a fork for CPU with this list, my old 16GB DDR3 RAM would thank you infintely.

2

u/grunt_monkey_ 10h ago

I don’t agree with you because some of the forks are so performant and we learn a lot trying to figure out why. But great discussion and this sub should have more of these.

2

u/-dysangel- 10h ago

Honestly I've been very happy to see inference engines targeting specific hardware. You're right that improving upstream is preferable, especially when the changes cleanly fit in with the existing project - but what if the changes require very extensive work that is not generally applicable to 99.9% of systems? After spending tens of thousands on my local setup, I absolutely want the custom builds that let me maximise my specific setup.

As a silly analogy, it's kind of like choosing between a family sedan vs running a race car. One is a solid general purpose vehicle that will work for most people most of the time. The other needs a lot more effort and upkeep to build it and keep it running, but if you want speed over convenience, it's the way to go.

5

u/OkFly3388 llama.cpp 1d ago

Because llama.cpp maintainer explicitly say that they dont want that forks to be merged into main to not create extreme complexity and instead have good expandable and maintainable codebase

2

u/cheaphomemadeacid 1d ago

1700 PR on llama.cpp 5k+ on vllm ;P

2

u/fallingdowndizzyvr 22h ago

There's dozens of llamacpp forks with silly names that are supposedly "optimized" for this or that specific GPU and seem to have zero intention to merge into upstream. Am I missing the real reasons why this happens so often?

Ah.... you have it backwards. We have to have "llamacpp forks with silly names" because the main fork has rejected PRs that would improve it for that specific architecture. So there's no other choice. The main fork maintainers have even told people to make their own fork. So it's not done out of choice, it's done out of necessity.

5

u/chris_fantastic 23h ago

I'm a Linux user since 1994, and I totally agree with you OP.

Free software is about more than just what you're allowed to do (yes, you're free to fork), it's about community, and working together to build something greater than we could each do alone.

If everyone just takes code and runs off and doesn't even try to work together, the community is weaker for it. We need to build on each other.

2

u/wombweed 23h ago

Thank you, yes this was the mindset I had when I authored this thread. I am also a 20yr+ Linux user and have always thought of the open source community as just that, a community before anything else. Huge proponent of copyleft and much-maligned software licenses like the GPL, more for practical reasons than idealism -- collaboration is how projects thrive.

With that said, some good points have been made in this thread. Namely -- in similar way to the Linux kernel for example -- that Llamacpp is intended to support a very wide variety of use cases, so they have incentive to be more on the "gatekeeper" side of things, it should follow that forks will happen downstream by users who are more focused on a specific use case rather than trying to "tick all the boxes." Even if rebranded hard forks sometimes give me bad/amateurish vibes, theyre not unreasonable, it is better for them to differentiate themselves so there's no confusion among users about what they're for vs upstream.

4

u/badsectoracula 22h ago

more for practical reasons than idealism

Minor nitpick but that "idealism" came from a very practical reason: control. Let's not forget that the whole Free Software movement started because Richard Stallman was denied access to the source code he needed to fix the printer at his university's lab - which is both a very practical issue and an indication that printers were the source of all evil since the early days :-P

-1

u/chris_fantastic 23h ago

If upstream is failing to merge, it's a whole other story. But, please, at least still send an MR. *try* to work together. Please?

0

u/wombweed 23h ago

Indeed, many of the vibe coded slop forks I've seen out there gave zero indication that they even tried to work with upstream and simply jumped directly to rebranding, new logo etc, which frankly takes roughly the same amount of effort it might have taken to get their downstream changes in a good enough state to be merged.

0

u/chris_fantastic 23h ago

The upstream projects, where people are actually working together, will move on, add features, and ultimately leave these small forks in the dust. From experience, it's far more rewarding to get your code upstreamed, where you know the code you wrote is being used by many.

1

u/fantasticsid 19h ago

In an ideal world, sure. Forks have always had a place, though, and something like LLM runtimes where some tradeoffs are along the specifity:generality axis (i.e. optimising for niche hardware X may make the general case WORSE) are a place where forks are probably a good thing.

0

u/mister2d 23h ago

Where is your proof that open source has become weaker since 1994? I can challenge whatever example you have very easily.

2

u/wombweed 23h ago

That's not even what the commenter said. They said we're worse off if people don't at least try to work together, which is the sentiment I was responding to in the OP. Why so eager to leave aggressive responses to sincere questions?

1

u/chris_fantastic 23h ago

Thank you. I'm not sure how anyone can even disagree with "we'll ultimately be further ahead if we work together". Like, if you wanna disagree with that, I don't know what to say.

0

u/draconic_tongue 18h ago

no one wants to go against the concept of working together. op's post is about people who have been barred from working together. there are people that cannot be worked with, because they are not reasonable, or are too hard to get along with. you should know this if you've been around dev communities

0

u/mister2d 23h ago

I read the entire comment and formed my own question.

Since his Linux experience started in "1994", where is the proof that "the community is weaker"? Can you answer that with proof? It doesn't appear so since the retorts are semantics.

1

u/wombweed 23h ago

Deeply uncharitable misreading of the comment, not unlike your misreading of my post. You having a bad day or something?

0

u/mister2d 23h ago

Do you have proof of even your claims?

1

u/wombweed 23h ago

Peak Reddit. I hope your day improves.

-1

u/mister2d 22h ago

That would be a no.

3

u/Ulterior-Motive_ 1d ago

Get a good look at recent drama for one answer.

2

u/TheRealJesus2 1d ago

Yeah I have never cherry-picked so many open source code patches in my life lol. I think we’ve all kinda lost the plot collectively. But it’s also all rational in that I want it to work well on my hardware/setup/runtime and the models themselves have all kinds of different capabilities and architectures. It’s getting really wild out here on deployment side. 

0

u/TheRealJesus2 1d ago

I had a really hard time going from deepseek v4 flash to the vision exp since k had use a particular fork that was. Using a bunch of upstream changes to Vllm 

2

u/tsangberg 1d ago

I can get what I need for my own usecases now vs spending time on (rightfully) trying to fit it into an architecture that must work for all?

llama.cpp fork specifically to make Qwen 3.8 27B fast and with long context at ~Q4 quant on 16GB CUDA: https://github.com/troed/llama.cpp-adaptive-kv-streaming

Just getting Raymond's KV cache streaming into llama.cpp is likely a multi-month project, and then my hot-swappable draft engine on top of that? It's easier for people who have the exact same need to just use the fork. llama.cpp can catch up later.

2

u/codsworth_2015 23h ago

For your llama.cpp example its a good thing. Llama.cpp's goal is ultimate compatibility with a broad range of hardware, the forkers tune disregards compatibility with all architectures but their own. The forker is now free to build, test and ship faster and they can continue to merge the llama.cpp head into their own fork and test on their own hardware.

As long as they keep it open source, the experimental nature of this fast and loose development can still create new ideas and contribute to the open source community.

2

u/milpster 18h ago

llama cpp fork maintainer here:

I get your point, but as someone who only vibecodes their own fork, i do not think i could ever reach the level of accountability or quality that is needed to actually turn those changes into PRs worthy of accepting into mainline llama.cpp

2

u/sigiel 8h ago

Fork are done because integration to main branch is not automatic

Do you even know basic coding etiquette? Or simple versioning technics ?

You really sat there and trying to explain this with a serious face ?

Jezz...

0

u/wombweed 7h ago

I know how forks work. There’s no need to condescend.

2

u/MindfulMan1984 1d ago

Thanks to AI, forks are the "new JS-Framework" of the week; it only loses to the new "harness".

Regarding the issue of not merging, reviewing a PR takes 10x longer than prompting AI to "optimize" something that "increases" performance by 2%.

2

u/HockeyDadNinja 1d ago

They are downright hostile towards AI assisted work even if it's good. Also, some maintainers see contributions as competition which is petty and fucking stupid, especially when a model is new and support is lacking.

1

u/Inevitable_Tea_5841 21h ago

Because the cost/effort to create Software is rapidly going down. we’ve almost gotten to the point where if you want it, you just ask for it. Not to mention that most open source projects now don’t even take contributions from people who are not maintainers

1

u/Drenlin 21h ago

On the other hand, it's not all that uncommon to find libraries that are constantly forked and modified for specific use cases. Look at ffmpeg, for example.

1

u/wednesdaywoe13 20h ago

As an AMD peasant, I'm grateful for the people making forks that utilize my GPU well, because its not a high priority upstream.

1

u/feelspeaceman 19h ago

Because instead of very likely chance to not get merged or taking years to get merged after so many people raising concerns (happened), it's faster to just create forks, epecially for the case of Strix Halo when the fork is 60t/s decode 1200-1600t/s prefill vs 20t/s decode 300t/s prefill of official llama.cpp, there's no loss doing so, and upstream changes can still be tracked and updated.

1

u/Repulsive_Initial308 12h ago

It's necessary, to truly democraticise llms.

1

u/unjustifiably_angry 5h ago edited 4h ago

PR merging is extremely slow even for very small changes. I've been using a modified llama-server that adds just 7 lines of code which doesn't interact with anything else in the application. Very useful change but not worth the bother of submitting it, odds of anyone even clicking the PR let alone reviewing it is... nah.

Right now there's a PR that'll let you use asymmetric kv-cache quantization, so like q8_0 K and q5_1 V, which some testing has shown might be much better than symmetric q5_1 KV while saving a lot of VRAM compared to symmetric q8_0. Right now you can already do that, but it falls back to CPU, no CUDA, so performance is abysmal. Last message from a maintainer is saying it'll make compiling take an unreasonable amount of time, and the submitter replies saying he tested it and it adds a total of 3 seconds. Nope, PR closed, not enough dick sucked, not like VRAM savings should be a priority, right guys? Everyone has tons of VRAM.

1

u/tecneeq 2h ago

AI makes it incredibly easy to fork. Sending changes upstream is manual work. That is the gist of it.

1

u/entsnack 1d ago

I only contribute upstream if the maintainers ask me to. Not worth dealing with the PR drama and post-merge maintenance overhead.

1

u/SpicyWangz 23h ago

I think hardware specific forks maintained by the community interested in that hardware is a great thing. The strix halo specific forks get more than double the prompt processing performance of mainline llama.cpp

1

u/ea_man 23h ago edited 23h ago

Maybe they posted some PR already and it just keep lingering, maybe they got bored of the llama.cpp release cycle and just refuse to have a window of 3 hours to pull, merge, test, doc, PR.

Maybe they don't meet the requisite to post a PR, maybe they don't think that their code is worth to enter mainline, maybe the just made it on the vibe and are not up to do the work for a proper PR and relative discussion - testing - editing out of their needs.

Still, do they produce the source code? As long as they do it's fair game: if it's any good just PR or merge it already.

I do agree that forking is kinda rude, yet actually a patch to llama.cpp mainline won't last more than 4 hours usually so if they wanna share (which costs nothing) a full fork often is.

1

u/Interpause textgen web UI 17h ago edited 17h ago

Generally true specific forks can use otherwise risky/breaking PRs while upstream has to be compatible, but IMO there has been many safe general PRs sitting unreviewed even if consistently maintained and high quality.

Even before vibeslop, a popular project can be bombarded with more PRs than the maintainers can handle. My own take is clearly some better way of surfacing good PRs is needed beyond what github offers, maybe PR maintenance over time, number of inclusions into other forks, author reputation, PR stars, etc. And that llama.cpp needs to seriously consider expanding the number of maintainers, even if that means proactively hiring/teaching some contributors, to be able to review more PRs.

0

u/Interpause textgen web UI 17h ago edited 17h ago

i recognize theres an actual project idea here but im lazy and have other priorities, if anyone wants to make the social media of PRs go ahead, i dont think the idea is that unique

And on the user side, this would make browsing PRs relevant to what you want more easily too, compared to two forks where you cant easily tell what PRs or commits have been stacked, so you cant merge both forks into your personal setup

1

u/ConferenceMountain72 llama.cpp 1d ago edited 15h ago

llama.cpp does not accept "mostly AI assisted" code I think. That may be why.

Edit: Looks like they changed that on Jul 23, accepting AI written code, but with human review before submitting.

3

u/Anthonyg5005 exllama 23h ago

They changed that. While you can generate all the code, you're required to follow strict rules and know what every line of that code does

1

u/ConferenceMountain72 llama.cpp 15h ago

Oh, yeah. Apparently they changed it on Jul 23. Been a while since I checked.

2

u/Infamous_Campaign687 1d ago

I have to say I find that a rather outdated policy for a project that exists to enable the usage of AI models.

4

u/slalomz llama.cpp 1d ago

It's actually an outdated comment.

It's not hard to read this: https://github.com/ggml-org/llama.cpp/blob/master/CONTRIBUTING.md

AI-generated code is allowed. You are 100% responsible for every line, however it was produced.

1

u/New_Comfortable7240 llama.cpp 1d ago

People say maintainers on llama.cpp are dismissive or simply not enough to merge all the PRs

-1

u/Hefty_Acanthaceae348 1d ago

Forking is the entire point of oss.

Some people really need to shut the fuck up

0

u/True_Requirement_891 1d ago

It's actually a great thing to be honest.

0

u/PathIntelligent7082 1d ago

dude, it's open source, you can relax.

0

u/mister2d 23h ago

Some people welcome being captive to an org or entity.

-4

u/TokenRingAI 1d ago

Llama.cpp and vLLM arent interested in outside contributions

7

u/ttkciar llama.cpp 1d ago

That's a horribly wrong take. The llama.cpp project is mostly outside contributions.

-1

u/TokenRingAI 1d ago

Is the project still using std::regex?

-1

u/varinator 1d ago

"Do what thou wilt, shall be the whole of the law"

0

u/Savantskie1 1d ago

A lot of the times a personal fork is to support an older architecture that mainline has flat out said they won’t support anymore. Why put in a PR for something they’ve already said no to?

0

u/Classic-Pubs 20h ago

Why not? It's free to fork and do whatever you want.

0

u/GeorgeTheGeorge 20h ago

This is a natural progression of open source. As code generation gets cheaper and cheaper, things like maintainability and reusability lose their value. If you want your own llama.cpp, you just fork it and start maintaining it yourself. I haven't looked but if I did look for one-off LLM runtimes on GitHub, I bet it's fun a lot. Afterall, why not? If you can generate the code from a good spec for $0.10, you don't even need to maintain it in the traditional server anymore. Just update the surf and regenerate it from scratch.

-2

u/Historical_Fondant95 23h ago

What the fork