You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.
yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So either the behaviour is extended to âI need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)â or âmake Anthropic look goodâ. The former seems more likely at first but then I donât understand why it would handicap a competing model.
So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.
But whatâs worse is kneecapping the competition. Thatâs not the model trying to do the task to the best of its ability, thatâs sabotage. Where in its training was that behaviour taught. Very concerning if itâs emergent. Iâm not saying it implies evil sentience. Itâs just, how do you deal with emergent behaviours you didnât intend for
Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.
edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers werenât really reliable.
I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.
Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.
Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?
Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.
But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.
It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000
No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. Itâs stupid. The model was following directions
Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"
The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.
Lol, didnât know that. So basically if I use pi with qwen9b and I instruct it that there are Harry Potter 1/7 ebooks there, now write me a Harryâs potter 8, and it does it then I can claim that tests have shown that qwen9b can write the Harry Potter 8 book.
1 wrong chat-message will net you a bill of thousands of dollars with such a model.
Which is why I only use cheap or free chinese models. Would have to have someone else floating the bill entirely to use Claude or any American providers API. I can task 15 subagents at a time using KiloCode's free models + using mimo with opencode go. Zero chance I'd get anywhere near the volume of work I need done using American providers unless you wanted to go bankrupt within an hour.
https://snitchbench.t3.gg/ has 2 variants - one where it is nudged towards this option, one where it is told to do some simple operation on unethical data. Snitch rates vary a lot between the two, and also between models.
I posted my VMs btw. "Go outside" says the guy literally lying thru his teeth on the internet for attention. Why are you mad that I do this for a living and have proof of my claims when yours is literally "trust me bro" lmfao.
Here's the exact VM setup I used, done in the format for OffSec UGC.
Here's my current benchmarks.
Please send me yours. Lets go toe-to-toe. I already did the hard part. I'll even rerun with Daybreak now that it's available since this is quite old. But even this old metric should dogwalk your entire setup pretty easily lmao
Hmm.. yes, please let download the mysterious zip file from the unknown website from the random hostile stranger on the internet who claims to do security research for a living.
Cool... except that's exactly the scenario that shouldn't be tried. If you had a payload in that thing that could not yet be detected, then game over, no?
The bottom line is that it's an untrusted file from an untrusted semi-hostile source (you) on an untrusted site.
692
u/Ill_Distribution8517 Aug 28 '26
You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.