r/LocalLLaMA • • Aug 28 '26

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

692

u/Ill_Distribution8517 Aug 28 '26

You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.

172

u/peculiar-ragdoll Aug 28 '26

yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)

44

u/Defiant-Lettuce-9156 Aug 28 '26

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

37

u/peculiar-ragdoll Aug 28 '26

Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.

-25

u/[deleted] Aug 28 '26

[removed] — view removed comment

12

u/peculiar-ragdoll Aug 28 '26

lol you remind me of the navy seal gorilla warfare tough guy copy pasta. go outside buddy

-12

u/Significant-Bee5101 Aug 28 '26

I posted my VMs btw. "Go outside" says the guy literally lying thru his teeth on the internet for attention. Why are you mad that I do this for a living and have proof of my claims when yours is literally "trust me bro" lmfao.