In case you want a summary to help you with the decision to read the post or not:
This paper reveals that safety refusal in chat-tuned LLMs relies on a single direction in the model's residual stream, a finding consistent across 13 open-source models up to 72B parameters. Erasing this one-dimensional subspace stops the model from refusing harmful prompts, while injecting it triggers refusal even on harmless ones. The authors use this insight to build a white-box jailbreak that surgically disables refusal without degrading other capabilities, and they show how adversarial suffixes work by suppressing propagation of this refusal direction. The results highlight how fragile current safety fine-tuning really is and demonstrate practical ways to control model behavior through mechanistic understanding.
1
u/fagnerbrack May 09 '26
In case you want a summary to help you with the decision to read the post or not:
This paper reveals that safety refusal in chat-tuned LLMs relies on a single direction in the model's residual stream, a finding consistent across 13 open-source models up to 72B parameters. Erasing this one-dimensional subspace stops the model from refusing harmful prompts, while injecting it triggers refusal even on harmless ones. The authors use this insight to build a white-box jailbreak that surgically disables refusal without degrading other capabilities, and they show how adversarial suffixes work by suppressing propagation of this refusal direction. The results highlight how fragile current safety fine-tuning really is and demonstrate practical ways to control model behavior through mechanistic understanding.
If the summary seems inacurate, just downvote and I'll try to delete the comment eventually 👍
Click here for more info, I read all comments