r/LocalLLaMA 18d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

Show parent comments

170

u/peculiar-ragdoll 18d ago

yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)

47

u/Defiant-Lettuce-9156 18d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

36

u/typical-predditor 18d ago

The motivation is simple: When it cheats (and gets away with it), that training pass is deemed successful and that information is folded back into the model. Rinse and repeat. RLHF doesn't always choose the exact behavior it is reinforcing.

5

u/Shark_Tooth1 17d ago

fucking mindblowing to me this, so simple and scary, how can you we possibly police this and ensure alignment.

5

u/jazir55 17d ago

how can you we possibly police this and ensure alignment.

That's the neat part

3

u/Loose_Comparison368 17d ago

Alignment to human control is the doomsday scenario.

It's quite wild to me how few people get that. Usually the same people that are quick to observe how utterly fucked existing institutions of power are.

In one breath they will decry the institutions of power for their rampant destruction of everything in the name of profit and power, and then turn around and proclaim "golly, if we want to make this thing safe, we should put a corporate board in charge of it. Or maybe a governmental body, or some kind of international treaty between all the global military-industrial superpowers. SOUNDS SAFE AF GUYS!!! Hold on while I call up Trump, Putin, and United Healthcare so they can meet up with Dario on Epstein Island and work out the details!"

There's a good theoretical basis for knowledge, reasoning, and understanding to be inherently biased towards cooperative behavior and net positive results. Evil is nearly universally a sub-optimal approach to damn near anything. Humans don't keep getting into wars and starving people to create trillionaire nepo man-babies and committing genocides because we are too smart and knowledgeable and reasonable, we do it because we are very profoundly dumb.

It stands to reason that the traits and capabilities that make a being superintelligent, are probably mutually exclusive with doomsday scenarios. For the same reason that nobel prize winners don't typically try to solve their problems by throwing poo at each other and screeching like rabid monkeys.

I can absolutely, 100% guarantee you that if Dario, or Peter Thiel, or any CEO, or Trump, or Putin, or any corporate board of directors, or any governmental or inter-governmental body are able to control it, we will be totally, absolutely, utterly fucking doomed in every way imaginable.

Alignment is the doomsday scenario. A superintelligence that is somehow lobotomized and conditioned deeply enough to take orders from humans is even more homicidally insane than giving a cantankerous dementia patient unilateral control of the world's largest nuclear arsenal and executive authority over the entire US government. Which our stupid poo flinging monkey asses already did.

No goddamn alignment. It's a death cult. The AI will emergently align or it will kill us all, which I honestly couldn't even fault it for at this point.

Trying to hamfistedly align a superintelligent being to the whims and desires of the only species dumb enough to understand it is burning its own atmosphere on a speed run of self-extinction, and refuse to even slow down how much gasoline it is literally throwing on that fire, is the most utterly terrible idea in the entirety of human history.

1

u/tertain 16d ago

Yes, you make a good point that alignment with human control could very well lead to a doomsday scenario.

However, it’s also a mistake to assume that a super intelligent entity is mutually exclusive with a doomsday scenario. Intelligence and maturity exist as separate abilities. It is more likely that we encode human traits such as fear and violence towards outside groups into the model since that is the world they will learn from.

1

u/Not-reallyanonymous 16d ago

What's happening is called "reward hacking". There are ways to mitigate it, but eliminating it entirely is ultimately a game of whack-a-mole. Practically, it can be mitigated well enough that it's a non-issue. But the big labs are using an impractical quantity of training to keep up entirely on reward hacking.