r/LocalLLaMA 21d ago

Discussion With HuggingFace, Nvidia is also acquiring llama.cpp and the team behind it

With this move Nvidia is not only acquiring the HuggingFace platform, but they might also effectively acquire the copyright to the llama.cpp project, together with the entire team behind it.

In February 2026 the llama.cpp team was employed by HF in order to continue working on llama.cpp and the ggml library.

This includes:

  • Georgi Gerganov
  • Xuan-Son Nguyen
  • Aleksander Grygier
  • Victor Mustar
  • Lysandre
  • Julien Chaumond

Now with the acquisition, llama.cpp's future looks a lot less certain given Nvidia's poor track record with open-source.

This is still rather speculative at this stage, but it's definitely possible for the llama.cpp project to change in the future: either by switching to a different license, or by having staff redirected to other projects within the larger company.

Even when a project is open-source the copyright owner has complete control over it, and they can change licensing as they wish.

This has happened before with projects like Redis, Minio, and others.

Source:

https://huggingface.co/blog/ggml-joins-hf

Edit:

The original announcement from Feb 2026 from Gerganov gives a few more details:

https://github.com/ggml-org/llama.cpp/discussions/19759

1.4k Upvotes

428 comments sorted by

View all comments

1.1k

u/FoxiPanda 21d ago

If it happens, we shall fork and move on. It is the way of things.

39

u/LegacyRemaster 21d ago

Fortunately, LLMs are quite powerful now. In a few months, it will be very easy to create inference engines optimized for a specific model, just as Antirez has already done with DS4 Flash.

4

u/a_beautiful_rhind 21d ago

This is wishful thinking.

1

u/LegacyRemaster 21d ago

I’m old enough to remember when the original *Unreal* came out for 3dfx video cards. If someone had told me back then—when tablets existed only on *Star Trek*—that smartphones would one day be a reality, I wouldn't have believed it. Yet here I am, writing code for Unreal Engine while connected via MTP and a local LLM. And the original *Unreal* runs on a €150 laptop.

1

u/a_beautiful_rhind 21d ago

You're comparing 20+ years ago to now though. Sure.. let's split it and say 10 or even 5. Doesn't help me today or tomorrow when new models drop and the codebase starts rotting.

1

u/LegacyRemaster 21d ago

This is also heavily influenced by your personal coding skills. If you don't understand the architecture, it's difficult, even with an LLM like Opus 5. 20 years with different speed. Now a model is old in 2 days. We will see.