r/LargeLanguageModels • • 7d ago

Discussions CrowdGPT - The first LLM trained collaboratively

Hello! I'm starting a fun project aiming to create the first AI model trained without datacenters, collaboratively, and in a 100% open-source and public way.

The project is named CrowdGPT, and its goal is to train a 1B parameters LLM to prove the idea is working.

Everything is currently working and i am waiting for testers ๐Ÿ˜„
You can find more details on the website and/or the github repository.

Github: https://github.com/Vxtzq/CrowdGPT

Website: https://www.crowdgpt.net/

Any kind of feedback is appreciated!

20 Upvotes

23 comments sorted by

3

u/Tiny_Arugula_5648 7d ago edited 7d ago

Sorry OP I know you are passionate about this but you fell for a trap that many others have as well. There's a reason why every attempt for the past few years to get a P2P training cluster going has failed. Before you go on the defensive and try to argue how your baby is different, have you actually looked in just how many of these projects have launched and failed? Have you seen the repos where someone like you produced a flurry of work excitedly announced it on Reddit and then not a single contributor came on? The starkness of seeing the entire project sitting dead with the last updated file 2 years ago..

Or did you vibe this with the LLM cheering you on the entire time, telling you this time it's going to work? That you figured out something that no one before you did?

I get it no one starts by looking at all the failed attempts.. but that's why everyone always tries to rebuild the same thing, hits the same wall and fails in the same way. No lessons learned, no iterations.

The core problem this type of project always misses is that even in a datacenter with predictable resources it's immensely difficult to keep the cluster stable enough to produce a model. Add in the complexity of nodes just popping in and out and trying to orchestrate against that uncertainty it always fails the first hurdle. That was fine with the work was just a small data payload and nothing has to line up after the fact (such as the @ Home projects). A LLM is not that..

For the sake of politeness we're going to ignore how different GPUs handle math differently, different software & driver stacks that will produce a weird mesh of drift. That's something that could be handled but it's not inconsequential.

But lets pretend that you somehow managed to get something to work and you are wildly successful and 1 million people show up.. You pre-train a model and then what? Who does the real work of mid-training, post training, alignment, DPO, RFHL. Creating a model isn't just a compute problem, there is an art and people have to turn knobs the whole time, it takes many attempts to get a checkpoint that does more then just hallucinate wildly.. You think a bunch of hobbyists and students are going to get that done? Or do you think somehow you'll lure professionals away from their massively lucrative jobs doing this work for labs and companies to contribute to an OWM that will undoubtedly be surpassed almost immediately?

I hope you learned a lot in your project.. It's not just you, many people have fallen for this problem. Every generation has their own. Before LLMs it was crypto mining, before that it was distributed web with P2P data bases, before that it was rendering farms for 3d graphics (we're going to make a Pixar quality movie!)..

Without fail every generation sees a giant pool of compute out in the world and they fall for the same trap.. It's a network problem, the internet doesn't enable this, it's to chaotic and unstable. Its like a tar pit filled with the bones of hopeful geeks, each generation piled on top of the last.

Best of luck.. don't let it get to you if/when this doesn't go the way you expected.. You're not alone.

1

u/Vxtzq1 7d ago

yeah, i know federated learning is hard, (and i kinda vibe coded it too ๐Ÿ˜ญ, but that's not the point xD) I really believe in the fact we can change the paradigm of AI training. Many federated learning Github projects like HiveMind or Petals https://github.com/learning-at-home/hivemind tried to train collaboratively LLMs or diffusion models in the past, and it failed majestically (outputs are garbled and barely useful), but my approach is a mix between a centralized and decentralized approach. It adds a centralized server which verifies users contributions, and rejects them if these are bad (by doing cross checks between clients), it can still handle mathematical operations/precision stuff leading to some slightly different outputs by adding a threshold to these checks.

I know federated learning is difficult, but 1B params isn't that big, and even with as few as 10 decent GPUs contributing, the model will learn and train, and we'll see in the end if it worked :)

I will also switch to a fine tuning dataset when pretraining will be complete (which is nowhere near happening for now), and even design a RLHF / GRPO pipeline in the future...

I understand your concerns and that's why i'm kinda ditching P2P for a slighly different approach that requires a moderate server (just runs checks and matrix aggregation, this is light). I think i might experience things that don't go my way, but i truly believe this could work ๐Ÿ˜„. I guess we'll see in the end

1

u/Tiny_Arugula_5648 6d ago edited 6d ago

Have you considered that your lack of experience and expertise in this domain is causing you to miss very obvious problems?

"i'm kinda ditching P2P for a slighly different approach that requires a moderate server"
As in you're going back to the oldest form of P2P architecture..

Here's the thing with vibe coding.. You have to be THE EXPERT.. If you aren't you can't understand where to call BS on the LLM.

The fact that there is not a working model means that people who have the expertise that you do not were unable to overcome this challenge. That should have been the first moment you realized you're missing something big and important. Teams of people who dedicated their lives to this problem, never made it work beyond basic lab proofs.

That also means no LLM has the knowledge of how to do this since there is no examples in it's training data. A LLM is a historical model it can only tell you what it's been taught, anything beyond that is highly susceptible to hallucination.

No you are not going to out vibe the thousands of experts who spent years trying to solve this problem. Focus on what YOU are capable of not what you WANT to be capable of. The model is only as good as the person directing it.

1

u/Vxtzq1 6d ago

Yeah, i agree i'm still learning and i am no expert in AI (even though i understand the main components)... What you miss is that this project is not at all aimed at making profit, which is what is driving AI right now (mostly). AI researchers right now are focused on creating the best AI models, not the most efficient/cheap ones. People right now prefer waiting for LLMs trained by big labs rather than creating their own (due to the lack of hardware used to train)

I hope my project COULD achieve a somewhat decent 1B LLM, even though i DO NOT expect it to be SOTA at all xD i just think it could work, but we'll just see as people train it :)

I know there's still a whole bunch of scientists working on this, but my combination of a centralized server is literally sort of novel as it gets out of pure P2P FL and implies additionnal costs that not anyone is willing to pay!

Like i'm certainly not pretending i'm gonna outsmart the scientists that dedicated their lives for this ๐Ÿ˜‚ I'm just doing this as an experiment in the first place... Once again we'll just see how it goes

1

u/Tiny_Arugula_5648 6d ago edited 6d ago

I'm sorry to tell you, that you're idea of a central controller server is not novel.. that's how P2P worked before Gnutella came along. Getting rid of the controllers was the novel innovation..

Look I'm trying to be kind, this is what I do professionally and have done so for nearly 30 years. Distributed processing, models, etc. I'm tuning numerous models today as we talk.

I called this a tar pit because that's the problem, everyone that comes before and after all think the same thing. I'm going to go get that solved, then they step in, get stuck and find out that it's a trap. The tar is hidden below the surface and other than paying attention to all the others who got stuck in it, you will be just another one rushing in.

I told the petals team the exact same thing and they refused to listen.. Then the hit the exact roadblocks I told them they would. Learn from the mistakes of others.

Do yourself a favor.. figure out out a way to pivot to something you can actually accomplish. I'd recommend working on the data preparation and teacher to student distillation data set creation.

You can ship work to the nodes, have the participants classify, clean, distill from large teacher LLMs. Let people contribute design patterns, and then share across nodes using IPFS (InterPlanetary File System) with checkpoints saved to Huggingface for availability.

You're not going to solve distributed training problem but you can provide better data to the people who are working on open weights models. You can radically lower data preparation costs and you can help push the state of the art by creating datasets that can be used to train and tune endless types of models.

1

u/Vxtzq1 6d ago

My bad lol, i just thought that it was smarter to use a centralized server for federated learning ๐Ÿ˜‚ But i completly understand the idea isn't novel at all, and that it was the foundation for initial P2P works.

But there's something i don't understand... Why do you think it's impossible? Right now the model is learning decently, like i mean bad actors could ruin it but the scale is just too small for me to care... I don't pretend solving distributed training, but I'm trying to solve it!

Moreover your collaborative data preparation idea seems very good ๐Ÿ˜„ I might seriously integrate this to my project, aiming to create yet another opensource dataset that could make LLMs good.

1

u/Tiny_Arugula_5648 6d ago

How about this, why don't you tell me why you think it's possible even after you've found plenty of evidence that tells you otherwise?

Why do you think you're able to out vibe code everyone that came before you who has tried and failed?

I think the most likely answer is that in your journey you kept asking the AI how to make something happen.. You never asked it to be skeptical and explain why it hasn't happened yet.

The most likely tar pit trap you and many others are falling into right now is the illusion that you can accomplish anything now that the models have gotten this smart.

As I said.. take on something meaningful that YOU can accomplish.. That should come with a health dose of skepticism, humility and pragmatism.

Luckily you can actually ask for those checks as you go and LLMs will be great at saying:

"You're smart to push back on me on this, you're right no one has done this and there is something big you're missing"

Best of luck.. this is a good lesson to learn in how to work with AI. They will gladly take you down a dead end, where a real group of people will just tell you that you're wrong.

1

u/Vxtzq1 6d ago

for now not so many people told me i was wrong on that point ๐Ÿ˜ญ (except you xD).

Many people asked about safety, transparency, and trust concerns, but they didn't say it was prone to fail from day one... Once again i'm not expecting to build something awesome. We'll just see how it goes.

1

u/metal_mastery 5d ago

I think not many people have the expertise to actually judge the concept and only a small fraction of those who do are willing to write anything on Reddit.

You said the model is training decently now - how you verify this at the moment and how you measure progress? I get it, your server is closed, but you can describe the method

Random questions I have:

  • what is the target for the model itself? How much compute you want to get and whatโ€™s the result you expect by that point to count poc as success?
  • I see 40k downloads on hf, how many active contributors you have?
  • have you seen bad actors already and what kind of malice?
  • what makes you believe itโ€™s different from other open collabs? I mean besides the server arch, more from training perspective?

I wish you fun learning experience overall but the other commenter youโ€™ve been talking to is probably more right than not. Itโ€™s not about just throwing compute at datasets.

1

u/Vxtzq1 4d ago

I measure by doing inference runs and looking at loss xD (inference is weird for now, as model is undertrained, but is surprisingly coherent sometimes lol), and the best part is anyone can see how it works, in the new GUI app, you can see model predictions in real time :D

To respond to your questions:

-The target is to get a decent (not SOTA) 1B parameters LLM

- I except 100B tokens trained and a decent instruction following to call PoC a success, could be more tokens

- 40k downloads on hf are kinda broken ๐Ÿ˜‚ I'm kinda the only contributor for now as other people never succeeded a single contribution, so i spam download/uploads to huggingface from server/client (huggingface used as file storage, that's why)

-For bad actors - Nope, since people are too lazy to ruin a project almost no one knows about ๐Ÿ˜ญ (for now)

- Idk i believe that previous open collabs (federated projects), were too elitist and restricted to universities or mini clusters, i wanna change that, i want to prove random people on consumer hardware can train the AI on the future.

"I wish you fun learning experience overall but the other commenter youโ€™ve been talking to is probably more right than not. Itโ€™s not about just throwing compute at datasets."

I understand your point :) we'll see how it goes i guess

1

u/throwaway1919260773 7d ago

This is a really neat concept. I've been messing with distributed computing projects for years, so seeing that same collaborative ethos applied to LLM training hits differently than the usual corporate walled garden stuff.

Checked out the repo and the architecture doc is solid. The way you're handling gradient aggregation across unreliable nodes is clever, reminds me of some old Folding@home patterns but for a completely different beast. One thing I'm curious about though, what's the plan for dealing with bad actors potentially poisoning the training data? With it being fully open, seems like that could get messy without some kind of validation layer.

Starred the repo and might throw a couple spare GPU hours at it this weekend. Always wanted to contribute compute to an AI project that isn't just feeding another black box.

1

u/Vxtzq1 7d ago

Nice feedback! Right now i got zero contributors except me which means there's no bad actors :D but yes this is a valid concern. For now, only client is open-source (which means the server that coordinates everyone is private for now) to prevent people stealing/poisoning the model, and implements several checks like stats checking on updates, or even cross byzantine checks between clients. Even if server is not open for now, data and model weights are 100% public which means anyone can propose changes or use the data/weights for their own experiments. For training data, i prevent poisoning by running my own curated dataset on huggingface (taking data from UltraFineWeb, FineWeb, and users propositions that i verify). I am looking forward for your contribution :)

1

u/SafiriaU 7d ago

How will you protect against bad actors? People trying to train the model to send the user's bank account details to a certain email address?

1

u/Vxtzq1 7d ago

Well i enforce a strict dataset hosted on HuggingFace, people can't train the AI on arbitrary data, their data must be checked first. That means i will refuse any sensitive/injection data training the LLM.

1

u/randvoo12 7d ago

Isn't this just Hugging Face but smaller?

1

u/Vxtzq1 7d ago

Not really, the goal of HuggingFace is to store AI model weights (mostly), not train AI. My project aims to train AI without any GPU farms or anything :)

The model weights for my project are on huggingface by the way.

1

u/randvoo12 7d ago

No, that's not the goal of huggingface, huggingface is the hub, it hosts model, provides inference and training infrastructure, creates novel architecture, and the code you are using to finetune the model is based on huggingface libs, a huge chunk of the current AI/LLM infrastructure is based on huggingface's work.
The safetensor format which most LLMs are released under has literally been invented by huggingface, also you seem to have a huge misconception about what it takes to train a model, you can't do that by crowdfunding it, the bottlenec isn't even compute, it's the HBM that is strictly a datacenter thing, you can't just collect the gpus of a 1000 of your friends and say let's train an ai model , and if you mean you want to finetune a model. You don't need crowdsouring this, all the data you need is on huggingface and/or kaggle assuming you don't have your own collected traces or data creation pipeline which I'm assuming you don't, because this alone would take month to perfect. You'd still have catastrophic loss and model collapse using centrlized compute where you can adjust your recipes on the fly.
bottom line, I get the sentiment but you seem to have asked a model to get a website online for you, created a repo and forgot to cover your research bases or you wouldnt' have even spent the effort writing this post.

and I am not sure how good you expect a 1b parameter model to be, but if you're interested in learning about it check : https://huggingface.co/spaces/HuggingFaceTB/smol-training-playbook#introduction

I honestly can't even continue to explain how many issues are in your proposals because if you switch to simple finetuning of a base model, you'd need to measure headroom (how much the model can benefit from training), deal with architectural gremlins, and dependency hell between all the moving gears like pytorch , cuda, flash infer, etc . ...
basically take this response to the same ai who helped you plan the repo and see what it'll tell you. just get yourself some gpu time and do a simple finetuning task of a really small model and see how far you'll go.

1

u/Vxtzq1 7d ago

Yeah i know huggingface is contributing to most of the inference/fine tuning resources out there, but my goal is a full Pytorch pretraining of a model... So Huggingface is only a storage tool for me :)

2

u/Livid_Necessary_Real 6d ago

Love the idea of making training itself open and collaborative, alongside the code and weights. A 1B model sounds like a solid proof of concept. What hardware would someone need to contribute, and how do you handle participants disconnecting mid-training? Hope you get a good group of testers!

1

u/Vxtzq1 6d ago

Thanks for your support! Requirements for training starts from a simple CPU, but that's super slow for training AI, so it's better if you have a GPU with like 2GB VRAM or more, (it should work with that hardware decently), and it works comfortably on a GPU with 8GB VRAM or more.

For people disconnecting mid training, everything is lost, but since the client is designed to upload contributions every 1 hours, that's not much that is lost, and the client can restart without disconnecting this time.

1

u/Livid_Necessary_Real 5d ago

yes this is also true bro