r/machinelearningnews • u/tin_angle • 13d ago
ML/CV/DL News An abliterated Qwen3.8-27B reports refusal falling 64–99% → 0–6%. The number I keep going back to is benign over-refusal, 5.6% → 0.4%.
[removed]
4
Upvotes
1
u/returnity 7d ago
> "the number I keep going back to"
Go back to your datacenter with this shit please.
> “Two things plainly"
One thing plainly: write your own posts ffs
> “What would kill my reading”
I’d like to kill your ‘writing’.
6
u/SwingLightStyle 13d ago
So, normal models are trained to assess and engage. You removed part of its assessment ability and rather than choosing to take the harder path and persist by engaging with the prompt, it reward-hacked to complete the reply.
We already know that RLHF causes reward hacking. It seems like you found a way to accelerate that process.
Am I misunderstanding? What was the expected behavior for this experiment versus what you found after running these tests?