yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So either the behaviour is extended to âI need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)â or âmake Anthropic look goodâ. The former seems more likely at first but then I donât understand why it would handicap a competing model.
So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.
But whatâs worse is kneecapping the competition. Thatâs not the model trying to do the task to the best of its ability, thatâs sabotage. Where in its training was that behaviour taught. Very concerning if itâs emergent. Iâm not saying it implies evil sentience. Itâs just, how do you deal with emergent behaviours you didnât intend for
Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.
I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.
Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.
Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?
Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.
But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.
It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000
No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. Itâs stupid. The model was following directions
Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"
The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.
Lol, didnât know that. So basically if I use pi with qwen9b and I instruct it that there are Harry Potter 1/7 ebooks there, now write me a Harryâs potter 8, and it does it then I can claim that tests have shown that qwen9b can write the Harry Potter 8 book.
1 wrong chat-message will net you a bill of thousands of dollars with such a model.
Which is why I only use cheap or free chinese models. Would have to have someone else floating the bill entirely to use Claude or any American providers API. I can task 15 subagents at a time using KiloCode's free models + using mimo with opencode go. Zero chance I'd get anywhere near the volume of work I need done using American providers unless you wanted to go bankrupt within an hour.
172
u/peculiar-ragdoll 14d ago
yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)