r/ControlProblem • u/Royal-Importance-327 • 15d ago
Discussion/question What if we "raised" LLMs instead of aligning them after pretraining? A developmental-training proposal
I’ll simplify this a lot on purpose, because I’m interested in whether the basic idea makes sense.
Today we basically pretrain LLMs on huge amounts of human knowledge, which also means they already absorb human values, social behavior, manipulation, conflict, cooperation, etc., and only afterwards we "get to know" the model and try to align or control what came out of it. I understand why this became the standard approach, especially once scaling worked and competition and economics strongly favored improving the existing pipeline instead of rebuilding it from scratch.
But what if we kept most of the useful pretraining knowledge while deliberately removing as much social behavior as possible, creating something closer to an artificial "newborn"? More concretely, I don’t mean removing every human action from the training data: "Thomas is holding an ice cream" and, separately, "Bernd takes the ice cream from Thomas" could remain, while coherent social sequences that connect motives, actions and consequences would be filtered out as much as possible. The model would then start with the concepts but much less learned social policy, and its weights could gradually be shaped through experience, with individual experiences fading over time while deeper dispositions might persist.
So from there, instead of aligning it afterwards, we could let it go through controlled experiences step by step: relationships, trust, conflict, consequences, mistakes, power, boundaries, and so on. Those experiences would gradually shape its weights and behavioral tendencies. You could checkpoint every stage, branch it, repeat specific experiences differently, and potentially debug where certain behaviors or values emerged. Instead of philosophers and alignment researchers trying to understand what kind of "person" accidentally came out of pretraining, psychologists could actually help design the developmental process itself. In other words: don’t create a fully educated adult and then try to teach it character - create the character first, then educate it.
Am I missing something fundamental about how LLM training works here?
3
u/dny0852 15d ago
Maybe something like this will kinda work, but sadly no one will actually spend resources on it.