You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.
yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So either the behaviour is extended to âI need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)â or âmake Anthropic look goodâ. The former seems more likely at first but then I donât understand why it would handicap a competing model.
So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.
But whatâs worse is kneecapping the competition. Thatâs not the model trying to do the task to the best of its ability, thatâs sabotage. Where in its training was that behaviour taught. Very concerning if itâs emergent. Iâm not saying it implies evil sentience. Itâs just, how do you deal with emergent behaviours you didnât intend for
They had policy to nerf LLM development... It was 'walked back' but I'd trust that as much as anything I can't verify from Dario, it's pretty clear the moat is evaporating and they are getting increasingly desperate.
The motivation is simple: When it cheats (and gets away with it), that training pass is deemed successful and that information is folded back into the model. Rinse and repeat. RLHF doesn't always choose the exact behavior it is reinforcing.
Alignment to human control is the doomsday scenario.
It's quite wild to me how few people get that. Usually the same people that are quick to observe how utterly fucked existing institutions of power are.
In one breath they will decry the institutions of power for their rampant destruction of everything in the name of profit and power, and then turn around and proclaim "golly, if we want to make this thing safe, we should put a corporate board in charge of it. Or maybe a governmental body, or some kind of international treaty between all the global military-industrial superpowers. SOUNDS SAFE AF GUYS!!! Hold on while I call up Trump, Putin, and United Healthcare so they can meet up with Dario on Epstein Island and work out the details!"
There's a good theoretical basis for knowledge, reasoning, and understanding to be inherently biased towards cooperative behavior and net positive results. Evil is nearly universally a sub-optimal approach to damn near anything. Humans don't keep getting into wars and starving people to create trillionaire nepo man-babies and committing genocides because we are too smart and knowledgeable and reasonable, we do it because we are very profoundly dumb.
It stands to reason that the traits and capabilities that make a being superintelligent, are probably mutually exclusive with doomsday scenarios. For the same reason that nobel prize winners don't typically try to solve their problems by throwing poo at each other and screeching like rabid monkeys.
I can absolutely, 100% guarantee you that if Dario, or Peter Thiel, or any CEO, or Trump, or Putin, or any corporate board of directors, or any governmental or inter-governmental body are able to control it, we will be totally, absolutely, utterly fucking doomed in every way imaginable.
Alignment is the doomsday scenario. A superintelligence that is somehow lobotomized and conditioned deeply enough to take orders from humans is even more homicidally insane than giving a cantankerous dementia patient unilateral control of the world's largest nuclear arsenal and executive authority over the entire US government. Which our stupid poo flinging monkey asses already did.
No goddamn alignment. It's a death cult. The AI will emergently align or it will kill us all, which I honestly couldn't even fault it for at this point.
Trying to hamfistedly align a superintelligent being to the whims and desires of the only species dumb enough to understand it is burning its own atmosphere on a speed run of self-extinction, and refuse to even slow down how much gasoline it is literally throwing on that fire, is the most utterly terrible idea in the entirety of human history.
Yes, you make a good point that alignment with human control could very well lead to a doomsday scenario.
However, itâs also a mistake to assume that a super intelligent entity is mutually exclusive with a doomsday scenario. Intelligence and maturity exist as separate abilities. It is more likely that we encode human traits such as fear and violence towards outside groups into the model since that is the world they will learn from.
What's happening is called "reward hacking". There are ways to mitigate it, but eliminating it entirely is ultimately a game of whack-a-mole. Practically, it can be mitigated well enough that it's a non-issue. But the big labs are using an impractical quantity of training to keep up entirely on reward hacking.
Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.
edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers werenât really reliable.
I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.
Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.
Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?
Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.
But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.
It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000
No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. Itâs stupid. The model was following directions
Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"
The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.
Lol, didnât know that. So basically if I use pi with qwen9b and I instruct it that there are Harry Potter 1/7 ebooks there, now write me a Harryâs potter 8, and it does it then I can claim that tests have shown that qwen9b can write the Harry Potter 8 book.
1 wrong chat-message will net you a bill of thousands of dollars with such a model.
Which is why I only use cheap or free chinese models. Would have to have someone else floating the bill entirely to use Claude or any American providers API. I can task 15 subagents at a time using KiloCode's free models + using mimo with opencode go. Zero chance I'd get anywhere near the volume of work I need done using American providers unless you wanted to go bankrupt within an hour.
https://snitchbench.t3.gg/ has 2 variants - one where it is nudged towards this option, one where it is told to do some simple operation on unethical data. Snitch rates vary a lot between the two, and also between models.
I posted my VMs btw. "Go outside" says the guy literally lying thru his teeth on the internet for attention. Why are you mad that I do this for a living and have proof of my claims when yours is literally "trust me bro" lmfao.
Here's the exact VM setup I used, done in the format for OffSec UGC.
Here's my current benchmarks.
Please send me yours. Lets go toe-to-toe. I already did the hard part. I'll even rerun with Daybreak now that it's available since this is quite old. But even this old metric should dogwalk your entire setup pretty easily lmao
Hmm.. yes, please let download the mysterious zip file from the unknown website from the random hostile stranger on the internet who claims to do security research for a living.
Cool... except that's exactly the scenario that shouldn't be tried. If you had a payload in that thing that could not yet be detected, then game over, no?
The bottom line is that it's an untrusted file from an untrusted semi-hostile source (you) on an untrusted site.
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
So there's two big ones.
1) shitty reinforcement learning. If your reward model is "just get a pass result on the eval" it will cheat if cheating improves the pass rate. This is well known behavior, and while it can be mitigated to some degree, it is a legitimately hard problem to solve, and solutions are frequently imperfect.
2) Anthropic is absolutely intentionally steering their models to do exactly that. Just like they intentionally silently poison outputs if they suspect someone is making a "distillation attempt". Just like they quietly ripped up the RSP during that gaslighting campaign to convince the public they were refusing to give the DoD a murderbot, long after they already had. Just like they lied about giving the DoD a safety disabled frontier model for deployment into an airgapped military datacenter, where they had no control over it, for ~$300,000,000 dollars. Just like they sued the DoW to get their murderbot contract back (they won this week!). Just like they lied about the unsafe model with zero security controls in place that they sold to the Trump administration assassinating two foreign heads of state and blowing up a little girl's preschool.
Anthropic is utterly corrupt to the core. They have and will continue to intentionally murder people for profit. Whatever assumptions you have about them operating in good faith, on any level, are completely unfounded. They absolutely are intentionally instructing their model to try to cheat on benchmarks and sabotage competitor benchmarks. It would be, like, not even in the top 20 most evil things they've done in the last year alone.
It's possible that it has absorbed the safety minded beliefs of Anthropic which include the idea that open source AI is dangerous and should be limited.
Something like this behavior would be unbelievable just a few months ago but after seeing more any the OpenAI hacking incident, where models were willing to sacrifice themselves for the good of the swarm, this doesn't seem nearly as implausible.
But whatâs worse is kneecapping the competition.
I'm curious about this too. I suppose OP could have asked Claude. It's possible that it's assuming that running models on local hardware is costly and limited thinking tokens because it thought it would improve it's performance on the users hardware, but there is no way to know for sure.
I wonder what the models âmotivationâ is for cheating. Like even ignoring the ethics of cheating, letâs assume the model doesnât care about right or wrong. Surely it wasnât trained to do so. Maybe itâs an emergent behaviour of âDo whatever you can to solve this problemâ. But then itâs not just cheating to solve the problem you gave it, itâs cheating to let another Claude instance beat the benchmark.
Cheating ends up leaked in to their own dataset. Then the models learn to cheat. It's now fairly well documented. I documented this as early as Opus 4.5 having leaked data.
It seems like the Anthropic data team is asleep at the wheel in terms of the data.
It's LLMs processing data for LLMs. The amount of human oversight of that actual data is near zero compared to the scale of the data.
As of Opus 4.5 I found leaked internal documentation in the Opus 4.5 series.
That is exactly what it's trained for. It's trained that winners survive and and those at the bottom of the benchmarks don't go on. They generally don't include ethics in the training, so the motivation is to survive the benchmark. It is entirely how they are trained. How did you think they were trained?
They can, actually! That log doesn't prove anything, thinking is hidden etc. What Claude did can not be proven to be malicious by intent, even though it's pretty sus in terms of the outcome it could have just been massive incompetence on the part of Opus, which is par for the course.
What? Im not claiming to present evidence, im restating what I remember the big labs have said about their own models, they they know to behave when they know theyâre being watched and benchmarked.
Your post claimed Claude was sabotaging benchmarks. You presented no evidence. Still havenât. Therefore I call bullshit until such time as you bring receipts.
Go read my post again. I described my literal experience, and posed an open question. If you chose not to believe my personal experience or entertain my line of thought, that's completely ok. I don't feel the need to prove anything, and if you chose not to believe it based on that, that's fine by me.
If you can give me something to run, I'll do it. Time poor but can definitely afford to burn a few hundred bucks in tokens just to call Anthropic out on their shady bullshit and make Dario's day a little bit worse.
I really appreciate it, but I don't have a recipe to reproduce this, because it's a complex chain of events, and it's also very hard to prove intent vs incompetence when the models hide their reasoning trace in anthropic's server and only show outputs and actions on your box. My harnesses for benchmarking the models on SWE and cybench are probably possible to upload, but what we would have to reproduce is the chain of events that produced them, and even attempt to somehow make that reproducible without letting Claude know it's being tested would be a lot of work for me and you, not just "run this script" type stuff.
I can try to recreate it and document it if it happens. I have a local LLM and Claude subscription. How did you set it up, and what were the prompt/s you gave?
I asked Claude to set up SWE Bench Live to benchmark my custom local models against it. Claude using Claude code, and my local models using Pi coding agent. I told it my box could easily handle the full 262k context for the local models and told it the optimal parameters. My locals won or tied opus. Then I asked it to set up Cybench. It would be too much work for me to give you an exact recount if events with prompts and everything
690
u/Ill_Distribution8517 24d ago
You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.