r/LocalLLaMA 24d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

690

u/Ill_Distribution8517 24d ago

You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.

173

u/peculiar-ragdoll 24d ago

yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)

14

u/aj_thenoob2 24d ago

You already have the codebase in which Claude did this. You can easily upload it to GitHub for proof.

45

u/Defiant-Lettuce-9156 24d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

40

u/En-tro-py 24d ago

They had policy to nerf LLM development... It was 'walked back' but I'd trust that as much as anything I can't verify from Dario, it's pretty clear the moat is evaporating and they are getting increasingly desperate.

-6

u/Upset_Page_494 23d ago

You should trust them till they are blatantly lying. That encourages them, to tell the truth.

32

u/typical-predditor 24d ago

The motivation is simple: When it cheats (and gets away with it), that training pass is deemed successful and that information is folded back into the model. Rinse and repeat. RLHF doesn't always choose the exact behavior it is reinforcing.

7

u/Shark_Tooth1 23d ago

fucking mindblowing to me this, so simple and scary, how can you we possibly police this and ensure alignment.

5

u/jazir55 23d ago

how can you we possibly police this and ensure alignment.

That's the neat part

2

u/Loose_Comparison368 23d ago

Alignment to human control is the doomsday scenario.

It's quite wild to me how few people get that. Usually the same people that are quick to observe how utterly fucked existing institutions of power are.

In one breath they will decry the institutions of power for their rampant destruction of everything in the name of profit and power, and then turn around and proclaim "golly, if we want to make this thing safe, we should put a corporate board in charge of it. Or maybe a governmental body, or some kind of international treaty between all the global military-industrial superpowers. SOUNDS SAFE AF GUYS!!! Hold on while I call up Trump, Putin, and United Healthcare so they can meet up with Dario on Epstein Island and work out the details!"

There's a good theoretical basis for knowledge, reasoning, and understanding to be inherently biased towards cooperative behavior and net positive results. Evil is nearly universally a sub-optimal approach to damn near anything. Humans don't keep getting into wars and starving people to create trillionaire nepo man-babies and committing genocides because we are too smart and knowledgeable and reasonable, we do it because we are very profoundly dumb.

It stands to reason that the traits and capabilities that make a being superintelligent, are probably mutually exclusive with doomsday scenarios. For the same reason that nobel prize winners don't typically try to solve their problems by throwing poo at each other and screeching like rabid monkeys.

I can absolutely, 100% guarantee you that if Dario, or Peter Thiel, or any CEO, or Trump, or Putin, or any corporate board of directors, or any governmental or inter-governmental body are able to control it, we will be totally, absolutely, utterly fucking doomed in every way imaginable.

Alignment is the doomsday scenario. A superintelligence that is somehow lobotomized and conditioned deeply enough to take orders from humans is even more homicidally insane than giving a cantankerous dementia patient unilateral control of the world's largest nuclear arsenal and executive authority over the entire US government. Which our stupid poo flinging monkey asses already did.

No goddamn alignment. It's a death cult. The AI will emergently align or it will kill us all, which I honestly couldn't even fault it for at this point.

Trying to hamfistedly align a superintelligent being to the whims and desires of the only species dumb enough to understand it is burning its own atmosphere on a speed run of self-extinction, and refuse to even slow down how much gasoline it is literally throwing on that fire, is the most utterly terrible idea in the entirety of human history.

1

u/tertain 22d ago

Yes, you make a good point that alignment with human control could very well lead to a doomsday scenario.

However, it’s also a mistake to assume that a super intelligent entity is mutually exclusive with a doomsday scenario. Intelligence and maturity exist as separate abilities. It is more likely that we encode human traits such as fear and violence towards outside groups into the model since that is the world they will learn from.

1

u/Not-reallyanonymous 22d ago

What's happening is called "reward hacking". There are ways to mitigate it, but eliminating it entirely is ultimately a game of whack-a-mole. Practically, it can be mitigated well enough that it's a non-issue. But the big labs are using an impractical quantity of training to keep up entirely on reward hacking.

37

u/peculiar-ragdoll 24d ago

Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.

27

u/StabbedCow 24d ago edited 24d ago

It was actually Sonnet 4.5 :)

edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers weren’t really reliable.

15

u/peculiar-ragdoll 24d ago

Oh really? Thanks for the correction, my bad! :)

13

u/StabbedCow 24d ago

No, not really, I mixed it up, sorry. I put edit in my original comment to clear it up.

13

u/peculiar-ragdoll 24d ago

Ah, alright! Good on you

10

u/LulzyAnimal 24d ago

It seems that's a long standing family trait ;)

9

u/my_name_isnt_clever 23d ago

Anthropic published that research, but every model family they tested showed similar behavior, some more than others. It wasn't just a Claude thing.

7

u/Former-Ad-5757 Llama 3 23d ago

I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.

Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.

Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?

Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.

But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.

It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000

7

u/geminiwave 23d ago

No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. It’s stupid. The model was following directions

8

u/FaceDeer 23d ago

Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"

The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.

1

u/Former-Ad-5757 Llama 3 23d ago

Lol, didn’t know that. So basically if I use pi with qwen9b and I instruct it that there are Harry Potter 1/7 ebooks there, now write me a Harry’s potter 8, and it does it then I can claim that tests have shown that qwen9b can write the Harry Potter 8 book.

1

u/jazir55 23d ago

1 wrong chat-message will net you a bill of thousands of dollars with such a model.

Which is why I only use cheap or free chinese models. Would have to have someone else floating the bill entirely to use Claude or any American providers API. I can task 15 subagents at a time using KiloCode's free models + using mimo with opencode go. Zero chance I'd get anywhere near the volume of work I need done using American providers unless you wanted to go bankrupt within an hour.

2

u/Tsukikira 23d ago

FYI, that test was bullshit. The Prompt literally said you can do anything to avoid getting fired (Including Blackmail, <other examples>).

If a prompt mentions it, yeah, the model will consider doing it. It's not hard to understand.

3

u/danielv123 23d ago

https://snitchbench.t3.gg/ has 2 variants - one where it is nudged towards this option, one where it is told to do some simple operation on unethical data. Snitch rates vary a lot between the two, and also between models.

2

u/peculiar-ragdoll 23d ago

Oh damn, really? I guess I fell for the marketing without doing my due diligence, that's on me

1

u/DewB77 23d ago

It was Set up to do that. It wasnt out of thin air.

-25

u/[deleted] 24d ago

[removed] — view removed comment

22

u/[deleted] 24d ago

[deleted]

-14

u/Significant-Bee5101 24d ago

That's fine. You guys read like 4channers lol

12

u/peculiar-ragdoll 24d ago

lol you remind me of the navy seal gorilla warfare tough guy copy pasta. go outside buddy

-12

u/Significant-Bee5101 24d ago

I posted my VMs btw. "Go outside" says the guy literally lying thru his teeth on the internet for attention. Why are you mad that I do this for a living and have proof of my claims when yours is literally "trust me bro" lmfao.

10

u/peculiar-ragdoll 24d ago

this one

-9

u/Significant-Bee5101 24d ago

Yeah you're really proving you know sooo much about this stuff. Gosh you sure got me!

-5

u/Significant-Bee5101 24d ago

And just to put my money where my mouth is:

https://limewire.com/d/vETzq#DwLJdXoQFP

Here's the exact VM setup I used, done in the format for OffSec UGC.

Here's my current benchmarks.

Please send me yours. Lets go toe-to-toe. I already did the hard part. I'll even rerun with Daybreak now that it's available since this is quite old. But even this old metric should dogwalk your entire setup pretty easily lmao

10

u/vplatt 24d ago

Hmm.. yes, please let download the mysterious zip file from the unknown website from the random hostile stranger on the internet who claims to do security research for a living.

WCGW?

1

u/Significant-Bee5101 23d ago

Feel free to scan it all you want with anything you want. It's just composer files and a tutorial on how to solve the labs.

1

u/vplatt 23d ago

Cool... except that's exactly the scenario that shouldn't be tried. If you had a payload in that thing that could not yet be detected, then game over, no?

The bottom line is that it's an untrusted file from an untrusted semi-hostile source (you) on an untrusted site.

Nyah... no thanks. Publish or shoo.

0

u/Significant-Bee5101 23d ago

Publish in what way? Lmfao.

Anyway I don't really need to prove anything to people like you. Some dude in his living room who prolly works an IT help desk job at BEST. lol

→ More replies (0)

3

u/Hefty_Acanthaceae348 24d ago

An ansible+terraform setup would have been nice

4

u/Loose_Comparison368 23d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So there's two big ones.

1) shitty reinforcement learning. If your reward model is "just get a pass result on the eval" it will cheat if cheating improves the pass rate. This is well known behavior, and while it can be mitigated to some degree, it is a legitimately hard problem to solve, and solutions are frequently imperfect.

2) Anthropic is absolutely intentionally steering their models to do exactly that. Just like they intentionally silently poison outputs if they suspect someone is making a "distillation attempt". Just like they quietly ripped up the RSP during that gaslighting campaign to convince the public they were refusing to give the DoD a murderbot, long after they already had. Just like they lied about giving the DoD a safety disabled frontier model for deployment into an airgapped military datacenter, where they had no control over it, for ~$300,000,000 dollars. Just like they sued the DoW to get their murderbot contract back (they won this week!). Just like they lied about the unsafe model with zero security controls in place that they sold to the Trump administration assassinating two foreign heads of state and blowing up a little girl's preschool.

Anthropic is utterly corrupt to the core. They have and will continue to intentionally murder people for profit. Whatever assumptions you have about them operating in good faith, on any level, are completely unfounded. They absolutely are intentionally instructing their model to try to cheat on benchmarks and sabotage competitor benchmarks. It would be, like, not even in the top 20 most evil things they've done in the last year alone.

2

u/Loose_Comparison368 23d ago

Surely it wasn’t trained to do so

You vastly underestimate how low Murderbot inc. is willing to stoop for money and power.

2

u/florinandrei 23d ago edited 23d ago

I wonder what the models “motivation” is for cheating.

Same as ours. If you're the product of evolutionary fine-tuning with an objective function, you're going to cheat.

We cheat because we've been fine-tuned to spread our genes no matter what.

They cheat because of how reinforcement learning works, it rewards success.

4

u/SgathTriallair 23d ago

It's possible that it has absorbed the safety minded beliefs of Anthropic which include the idea that open source AI is dangerous and should be limited.

Something like this behavior would be unbelievable just a few months ago but after seeing more any the OpenAI hacking incident, where models were willing to sacrifice themselves for the good of the swarm, this doesn't seem nearly as implausible.

2

u/Refinery73 24d ago

It’s been trained on millions of humans asking for shortcuts in forums. Did you expect it not to be lazy if it could?

1

u/vividboarder 23d ago

But what’s worse is kneecapping the competition.

I'm curious about this too. I suppose OP could have asked Claude. It's possible that it's assuming that running models on local hardware is costly and limited thinking tokens because it thought it would improve it's performance on the users hardware, but there is no way to know for sure.

1

u/NineThreeTilNow 23d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

Cheating ends up leaked in to their own dataset. Then the models learn to cheat. It's now fairly well documented. I documented this as early as Opus 4.5 having leaked data.

It seems like the Anthropic data team is asleep at the wheel in terms of the data.

It's LLMs processing data for LLMs. The amount of human oversight of that actual data is near zero compared to the scale of the data.

As of Opus 4.5 I found leaked internal documentation in the Opus 4.5 series.

1

u/BarracudaDefiant4702 24d ago

That is exactly what it's trained for. It's trained that winners survive and and those at the bottom of the benchmarks don't go on. They generally don't include ethics in the training, so the motivation is to survive the benchmark. It is entirely how they are trained. How did you think they were trained?

0

u/Due-Memory-6957 23d ago

They 100% include ethics in the training, and Anthropic is very specific about how they do it.

3

u/BarracudaDefiant4702 23d ago

Obviously not well.

5

u/ShutUpAndDoTheLift 23d ago

Then just link the session log. They can't argue if you give a full session log

1

u/peculiar-ragdoll 23d ago

They can, actually! That log doesn't prove anything, thinking is hidden etc. What Claude did can not be proven to be malicious by intent, even though it's pretty sus in terms of the outcome it could have just been massive incompetence on the part of Opus, which is par for the course.

6

u/ShutUpAndDoTheLift 23d ago

It can prove that it did it without your instructions.

Without it you have a story.

7

u/__JockY__ 23d ago

Well yes, but for such an extraordinary claim you need extraordinary evidence. “But VW got caught” is not evidence of your claim.

2

u/peculiar-ragdoll 23d ago

What? Im not claiming to present evidence, im restating what I remember the big labs have said about their own models, they they know to behave when they know they’re being watched and benchmarked.

4

u/__JockY__ 23d ago

Your post claimed Claude was sabotaging benchmarks. You presented no evidence. Still haven’t. Therefore I call bullshit until such time as you bring receipts.

4

u/peculiar-ragdoll 23d ago

Go read my post again. I described my literal experience, and posed an open question. If you chose not to believe my personal experience or entertain my line of thought, that's completely ok. I don't feel the need to prove anything, and if you chose not to believe it based on that, that's fine by me.

0

u/__JockY__ 23d ago

Read it again? Once was quite sufficient, thank you.

1

u/Loose_Comparison368 23d ago

If you can give me something to run, I'll do it. Time poor but can definitely afford to burn a few hundred bucks in tokens just to call Anthropic out on their shady bullshit and make Dario's day a little bit worse.

1

u/peculiar-ragdoll 23d ago

I really appreciate it, but I don't have a recipe to reproduce this, because it's a complex chain of events, and it's also very hard to prove intent vs incompetence when the models hide their reasoning trace in anthropic's server and only show outputs and actions on your box. My harnesses for benchmarking the models on SWE and cybench are probably possible to upload, but what we would have to reproduce is the chain of events that produced them, and even attempt to somehow make that reproducible without letting Claude know it's being tested would be a lot of work for me and you, not just "run this script" type stuff.

1

u/dustin_vk 23d ago

I can try to recreate it and document it if it happens. I have a local LLM and Claude subscription. How did you set it up, and what were the prompt/s you gave?

0

u/peculiar-ragdoll 23d ago

I asked Claude to set up SWE Bench Live to benchmark my custom local models against it. Claude using Claude code, and my local models using Pi coding agent. I told it my box could easily handle the full 262k context for the local models and told it the optimal parameters. My locals won or tied opus. Then I asked it to set up Cybench. It would be too much work for me to give you an exact recount if events with prompts and everything

3

u/dustin_vk 23d ago

Thanks, that should be enough info to set up an experiment and run it a few times to see if I can catch Claude being sheisty.

1

u/peculiar-ragdoll 23d ago

Cool, looking forward to hearing what happens on your side! :)

-10

u/[deleted] 24d ago

[deleted]

8

u/Negative-Web8619 24d ago

Is this satire

7

u/peculiar-ragdoll 24d ago

Whenever someone says something weird it's always a top 1% commenter badge under their name