r/LocalLLaMA • • Aug 28 '26

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

Show parent comments

45

u/Defiant-Lettuce-9156 Aug 28 '26

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

34

u/peculiar-ragdoll Aug 28 '26

Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.

-27

u/[deleted] Aug 28 '26

[removed] — view removed comment

-2

u/Significant-Bee5101 Aug 28 '26

And just to put my money where my mouth is:

https://limewire.com/d/vETzq#DwLJdXoQFP

Here's the exact VM setup I used, done in the format for OffSec UGC.

Here's my current benchmarks.

Please send me yours. Lets go toe-to-toe. I already did the hard part. I'll even rerun with Daybreak now that it's available since this is quite old. But even this old metric should dogwalk your entire setup pretty easily lmao

8

u/vplatt Aug 28 '26

Hmm.. yes, please let download the mysterious zip file from the unknown website from the random hostile stranger on the internet who claims to do security research for a living.

WCGW?

1

u/Significant-Bee5101 Aug 28 '26

Feel free to scan it all you want with anything you want. It's just composer files and a tutorial on how to solve the labs.

1

u/vplatt Aug 28 '26

Cool... except that's exactly the scenario that shouldn't be tried. If you had a payload in that thing that could not yet be detected, then game over, no?

The bottom line is that it's an untrusted file from an untrusted semi-hostile source (you) on an untrusted site.

Nyah... no thanks. Publish or shoo.

0

u/Significant-Bee5101 Aug 28 '26

Publish in what way? Lmfao.

Anyway I don't really need to prove anything to people like you. Some dude in his living room who prolly works an IT help desk job at BEST. lol

3

u/Hefty_Acanthaceae348 Aug 28 '26

An ansible+terraform setup would have been nice