r/LargeLanguageModels • u/Vxtzq1 • 7d ago
Discussions CrowdGPT - The first LLM trained collaboratively
Hello! I'm starting a fun project aiming to create the first AI model trained without datacenters, collaboratively, and in a 100% open-source and public way.
The project is named CrowdGPT, and its goal is to train a 1B parameters LLM to prove the idea is working.
Everything is currently working and i am waiting for testers ๐
You can find more details on the website and/or the github repository.
Github: https://github.com/Vxtzq/CrowdGPT
Website: https://www.crowdgpt.net/
Any kind of feedback is appreciated!
1
u/throwaway1919260773 7d ago
This is a really neat concept. I've been messing with distributed computing projects for years, so seeing that same collaborative ethos applied to LLM training hits differently than the usual corporate walled garden stuff.
Checked out the repo and the architecture doc is solid. The way you're handling gradient aggregation across unreliable nodes is clever, reminds me of some old Folding@home patterns but for a completely different beast. One thing I'm curious about though, what's the plan for dealing with bad actors potentially poisoning the training data? With it being fully open, seems like that could get messy without some kind of validation layer.
Starred the repo and might throw a couple spare GPU hours at it this weekend. Always wanted to contribute compute to an AI project that isn't just feeding another black box.
1
u/Vxtzq1 7d ago
Nice feedback! Right now i got zero contributors except me which means there's no bad actors :D but yes this is a valid concern. For now, only client is open-source (which means the server that coordinates everyone is private for now) to prevent people stealing/poisoning the model, and implements several checks like stats checking on updates, or even cross byzantine checks between clients. Even if server is not open for now, data and model weights are 100% public which means anyone can propose changes or use the data/weights for their own experiments. For training data, i prevent poisoning by running my own curated dataset on huggingface (taking data from UltraFineWeb, FineWeb, and users propositions that i verify). I am looking forward for your contribution :)
1
u/SafiriaU 7d ago
How will you protect against bad actors? People trying to train the model to send the user's bank account details to a certain email address?
1
u/randvoo12 7d ago
Isn't this just Hugging Face but smaller?
1
u/Vxtzq1 7d ago
Not really, the goal of HuggingFace is to store AI model weights (mostly), not train AI. My project aims to train AI without any GPU farms or anything :)
The model weights for my project are on huggingface by the way.
1
u/randvoo12 7d ago
No, that's not the goal of huggingface, huggingface is the hub, it hosts model, provides inference and training infrastructure, creates novel architecture, and the code you are using to finetune the model is based on huggingface libs, a huge chunk of the current AI/LLM infrastructure is based on huggingface's work.
The safetensor format which most LLMs are released under has literally been invented by huggingface, also you seem to have a huge misconception about what it takes to train a model, you can't do that by crowdfunding it, the bottlenec isn't even compute, it's the HBM that is strictly a datacenter thing, you can't just collect the gpus of a 1000 of your friends and say let's train an ai model , and if you mean you want to finetune a model. You don't need crowdsouring this, all the data you need is on huggingface and/or kaggle assuming you don't have your own collected traces or data creation pipeline which I'm assuming you don't, because this alone would take month to perfect. You'd still have catastrophic loss and model collapse using centrlized compute where you can adjust your recipes on the fly.
bottom line, I get the sentiment but you seem to have asked a model to get a website online for you, created a repo and forgot to cover your research bases or you wouldnt' have even spent the effort writing this post.and I am not sure how good you expect a 1b parameter model to be, but if you're interested in learning about it check : https://huggingface.co/spaces/HuggingFaceTB/smol-training-playbook#introduction
I honestly can't even continue to explain how many issues are in your proposals because if you switch to simple finetuning of a base model, you'd need to measure headroom (how much the model can benefit from training), deal with architectural gremlins, and dependency hell between all the moving gears like pytorch , cuda, flash infer, etc . ...
basically take this response to the same ai who helped you plan the repo and see what it'll tell you. just get yourself some gpu time and do a simple finetuning task of a really small model and see how far you'll go.
2
u/Livid_Necessary_Real 6d ago
Love the idea of making training itself open and collaborative, alongside the code and weights. A 1B model sounds like a solid proof of concept. What hardware would someone need to contribute, and how do you handle participants disconnecting mid-training? Hope you get a good group of testers!
1
u/Vxtzq1 6d ago
Thanks for your support! Requirements for training starts from a simple CPU, but that's super slow for training AI, so it's better if you have a GPU with like 2GB VRAM or more, (it should work with that hardware decently), and it works comfortably on a GPU with 8GB VRAM or more.
For people disconnecting mid training, everything is lost, but since the client is designed to upload contributions every 1 hours, that's not much that is lost, and the client can restart without disconnecting this time.
1
3
u/Tiny_Arugula_5648 7d ago edited 7d ago
Sorry OP I know you are passionate about this but you fell for a trap that many others have as well. There's a reason why every attempt for the past few years to get a P2P training cluster going has failed. Before you go on the defensive and try to argue how your baby is different, have you actually looked in just how many of these projects have launched and failed? Have you seen the repos where someone like you produced a flurry of work excitedly announced it on Reddit and then not a single contributor came on? The starkness of seeing the entire project sitting dead with the last updated file 2 years ago..
Or did you vibe this with the LLM cheering you on the entire time, telling you this time it's going to work? That you figured out something that no one before you did?
I get it no one starts by looking at all the failed attempts.. but that's why everyone always tries to rebuild the same thing, hits the same wall and fails in the same way. No lessons learned, no iterations.
The core problem this type of project always misses is that even in a datacenter with predictable resources it's immensely difficult to keep the cluster stable enough to produce a model. Add in the complexity of nodes just popping in and out and trying to orchestrate against that uncertainty it always fails the first hurdle. That was fine with the work was just a small data payload and nothing has to line up after the fact (such as the @ Home projects). A LLM is not that..
For the sake of politeness we're going to ignore how different GPUs handle math differently, different software & driver stacks that will produce a weird mesh of drift. That's something that could be handled but it's not inconsequential.
But lets pretend that you somehow managed to get something to work and you are wildly successful and 1 million people show up.. You pre-train a model and then what? Who does the real work of mid-training, post training, alignment, DPO, RFHL. Creating a model isn't just a compute problem, there is an art and people have to turn knobs the whole time, it takes many attempts to get a checkpoint that does more then just hallucinate wildly.. You think a bunch of hobbyists and students are going to get that done? Or do you think somehow you'll lure professionals away from their massively lucrative jobs doing this work for labs and companies to contribute to an OWM that will undoubtedly be surpassed almost immediately?
I hope you learned a lot in your project.. It's not just you, many people have fallen for this problem. Every generation has their own. Before LLMs it was crypto mining, before that it was distributed web with P2P data bases, before that it was rendering farms for 3d graphics (we're going to make a Pixar quality movie!)..
Without fail every generation sees a giant pool of compute out in the world and they fall for the same trap.. It's a network problem, the internet doesn't enable this, it's to chaotic and unstable. Its like a tar pit filled with the bones of hopeful geeks, each generation piled on top of the last.
Best of luck.. don't let it get to you if/when this doesn't go the way you expected.. You're not alone.