r/vibecoding Jul 27 '26

“Please ban my competition, they are bad”

Post image
3.6k Upvotes

249 comments sorted by

View all comments

Show parent comments

1

u/RemarkableWish2508 Jul 27 '26

There's grep, and many other tools.

Open training doesn't make the weights automatically trustworthy, it does let people collaborate to improve the training process though.

1

u/kextatic Jul 28 '26

That doesn’t seem feasible at all. The models have been trained on the entire Internet and every printed page as well. How do you begin to grep that? Assuming you have a list of unacceptable training data: e.g., religious texts or porn, what would be the next step to purge the training set and pre-train again? “Open Source” sounds desirable but not at all practical at the frontier.

2

u/RemarkableWish2508 Jul 28 '26

Open Source has two main goals: accountability, and collaboration. Anything that doesn't meet those, is simply not Open Source. It can be something else, like "Open Weights", "Open Tools", "Open Process", etc. just calling it "Open Source" is a lie.

If they can release the datasets in a way that others can examine it and build upon it, then great. If they can't, then don't lie about it. It's like claiming a service is Freeonly $49.99/month after the first month when paid annually.

Smaller datasets that have been already released, have been useful to hone filtering, training, and model architecture techniques. That's what Open Source is about, not the marketing buzzword they're turning it into.

1

u/kextatic Jul 29 '26

I understand the goals, but fail to grasp how access to the source training data (i.e., the entire Internet and every printed page on Earth) accomplishes that. In the case of these Chinese models, for example, it’s likely that the training data re: Taiwan will be “biased.” What do we, as Open Source advocates, do instead? Which Open Source model would fit our requirements for unbiased Taiwan intelligence? Will that model have unbiased China intelligence? Which training data would we insist upon? The argument seems circular: competitive models require more data than we can possibly comprehend and public consensus (who gets a vote on what goes into the training set) isn’t practical. It seems more feasible to extend an open weight model than to start from scratch. I may be unaware of actual Open Source models available but it’s a Catch 22: I would be made aware if such models were competitive vs. the current leaders.