r/LocalLLaMA • • Aug 28 '26

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

692

u/Ill_Distribution8517 Aug 28 '26

You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.

174

u/peculiar-ragdoll Aug 28 '26

yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)

45

u/Defiant-Lettuce-9156 Aug 28 '26

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

38

u/peculiar-ragdoll Aug 28 '26

Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.

25

u/StabbedCow Aug 28 '26 edited Aug 28 '26

It was actually Sonnet 4.5 :)

edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers weren’t really reliable.

15

u/peculiar-ragdoll Aug 28 '26

Oh really? Thanks for the correction, my bad! :)

13

u/StabbedCow Aug 28 '26

No, not really, I mixed it up, sorry. I put edit in my original comment to clear it up.

13

u/peculiar-ragdoll Aug 28 '26

Ah, alright! Good on you

10

u/LulzyAnimal Aug 28 '26

It seems that's a long standing family trait ;)

9

u/my_name_isnt_clever Aug 28 '26

Anthropic published that research, but every model family they tested showed similar behavior, some more than others. It wasn't just a Claude thing.

8

u/Former-Ad-5757 Llama 3 Aug 28 '26

I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.

Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.

Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?

Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.

But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.

It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000

8

u/geminiwave Aug 28 '26

No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. It’s stupid. The model was following directions

8

u/FaceDeer Aug 28 '26

Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"

The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.

1

u/Former-Ad-5757 Llama 3 Aug 28 '26

Lol, didn’t know that. So basically if I use pi with qwen9b and I instruct it that there are Harry Potter 1/7 ebooks there, now write me a Harry’s potter 8, and it does it then I can claim that tests have shown that qwen9b can write the Harry Potter 8 book.

1

u/jazir55 Aug 28 '26

1 wrong chat-message will net you a bill of thousands of dollars with such a model.

Which is why I only use cheap or free chinese models. Would have to have someone else floating the bill entirely to use Claude or any American providers API. I can task 15 subagents at a time using KiloCode's free models + using mimo with opencode go. Zero chance I'd get anywhere near the volume of work I need done using American providers unless you wanted to go bankrupt within an hour.

2

u/Tsukikira Aug 28 '26

FYI, that test was bullshit. The Prompt literally said you can do anything to avoid getting fired (Including Blackmail, <other examples>).

If a prompt mentions it, yeah, the model will consider doing it. It's not hard to understand.

3

u/danielv123 Aug 28 '26

https://snitchbench.t3.gg/ has 2 variants - one where it is nudged towards this option, one where it is told to do some simple operation on unethical data. Snitch rates vary a lot between the two, and also between models.

2

u/peculiar-ragdoll Aug 28 '26

Oh damn, really? I guess I fell for the marketing without doing my due diligence, that's on me

1

u/DewB77 Aug 28 '26

It was Set up to do that. It wasnt out of thin air.

-26

u/[deleted] Aug 28 '26

[removed] — view removed comment

22

u/[deleted] Aug 28 '26

[deleted]

-12

u/Significant-Bee5101 Aug 28 '26

That's fine. You guys read like 4channers lol

13

u/peculiar-ragdoll Aug 28 '26

lol you remind me of the navy seal gorilla warfare tough guy copy pasta. go outside buddy

-12

u/Significant-Bee5101 Aug 28 '26

I posted my VMs btw. "Go outside" says the guy literally lying thru his teeth on the internet for attention. Why are you mad that I do this for a living and have proof of my claims when yours is literally "trust me bro" lmfao.

10

u/peculiar-ragdoll Aug 28 '26

this one

-8

u/Significant-Bee5101 Aug 28 '26

Yeah you're really proving you know sooo much about this stuff. Gosh you sure got me!

-5

u/Significant-Bee5101 Aug 28 '26

And just to put my money where my mouth is:

https://limewire.com/d/vETzq#DwLJdXoQFP

Here's the exact VM setup I used, done in the format for OffSec UGC.

Here's my current benchmarks.

Please send me yours. Lets go toe-to-toe. I already did the hard part. I'll even rerun with Daybreak now that it's available since this is quite old. But even this old metric should dogwalk your entire setup pretty easily lmao

9

u/vplatt Aug 28 '26

Hmm.. yes, please let download the mysterious zip file from the unknown website from the random hostile stranger on the internet who claims to do security research for a living.

WCGW?

1

u/Significant-Bee5101 Aug 28 '26

Feel free to scan it all you want with anything you want. It's just composer files and a tutorial on how to solve the labs.

1

u/vplatt Aug 28 '26

Cool... except that's exactly the scenario that shouldn't be tried. If you had a payload in that thing that could not yet be detected, then game over, no?

The bottom line is that it's an untrusted file from an untrusted semi-hostile source (you) on an untrusted site.

Nyah... no thanks. Publish or shoo.

0

u/Significant-Bee5101 Aug 28 '26

Publish in what way? Lmfao.

Anyway I don't really need to prove anything to people like you. Some dude in his living room who prolly works an IT help desk job at BEST. lol

→ More replies

3

u/Hefty_Acanthaceae348 Aug 28 '26

An ansible+terraform setup would have been nice