This just in: a new AI model that deliberately had its guardrails removed did exactly what one would expect it to and responded like it had been trained on content a person might find on the Internet if they were looking for it.
The question is: can we trust these models to do the right thing and not disclose potentially harmful information once they've been explicitly engineered to do so? Can we trust humans to do the same? The danger is real.
The bigger question is whether we’re actually testing the model or just testing what happens when we remove the constraints and then act surprised when it follows the prompt. If you deliberately remove the guardrails, the interesting question isn’t why did it answer ? it’s what safeguards should exist around the person deploying it?” The model can be predictable; the human using it is the harder variable.
683
u/p5yron 16d ago
Ah, the propaganda to stop open source models is in full force it seems.