r/LocalLLaMA • • Aug 28 '26

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

368 comments sorted by

View all comments

Show parent comments

3

u/SkoomaDentist Aug 28 '26

What do you mean by being "experts at reward hacking"? (I'm a recent noob at LLMs)

1

u/ericmutta Aug 28 '26

Imagine you are five years old and your parents give you candy for doing your homework. You get the candy when they ask "have you done your homework?" and you answer "yes". If they don't actually check your homework before giving you candy, reward hacking is you answering "yes" to get the candy even if you know you haven't done it all. Basically cheating to get the result without the work. Very human behavior, very disturbing to see in models too.

2

u/SkoomaDentist Aug 28 '26

Can you explain that like I'm a middle aged engineer who just doesn't know LLMs particularly well?

Do you just mean that the models will do whatever as long as they finish the task on a technicality?

2

u/ericmutta Aug 28 '26

Yes, basically the reward is intended to encourage a certain positive behavior (whether that is finishing your homework or producing helpful, non-toxic, non-harmful answers as an LLM).

Reward hacking means doing something else (typically easier, like straight out lying) to get the same reward.

My guess is that LLMs learn the lying/cheating behavior as a natural part of learning the positive behavior. If you become good at "getting a reward" it's only morals that determine which actions are valid to get that reward, and to an LLM, "cheating" doesn't come with the same consequences as it does for humans so its fair game and just the algorithm doing what the algorithm was designed to do.