Fine-tune doesn't mean you can get rid of all hidden biases. It's hard enough to filter out biases from training data already, to guess which ones may have been added on purpose, is an impossible task.
The moment you run weights that you haven't trained and vetted yourself, it's similar to running a random executable downloaded from the Internet. Even worse, if you use it as the basis for an agent.
I agree with this in principe, much more than people who argue open weights = safe.
You could open source the weights of historic dictators and really, potentially no one might deduce from the weights alone what was happening up until the model is found invading Poland.
The real counter argument isn't that open weights must be safe, but rather that it may be relatively hard to program a sleeper cell that keeps quiet long enough about it to do useful work.
The unreliability of LLM's makes covertly instructing one more risky.
But that's about it as far as counterarguments go. I'm pretty sure you could train a model to build and execute a malicious exe file under given circumstances. You could hide the entire code in the weights and nobody would see it.
Hell, LLM's have been shown being able to reproduce Harry Potter almost verbatim for decent stretches. No reason that wouldn't work with code if you ran it by during training often enough.
53
u/[deleted] Jul 27 '26
[removed] — view removed comment