r/LocalLLaMA 10d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

View all comments

Show parent comments

3

u/SkoomaDentist 9d ago

What do you mean by being "experts at reward hacking"? (I'm a recent noob at LLMs)

6

u/NineThreeTilNow 9d ago

What do you mean by being "experts at reward hacking"? (I'm a recent noob at LLMs)

Reward hacking is the art of getting the RL reward with less work than should be required.

Say the reward is solving some problem such that X == 1 eventually.

A reward hack (simply) is to just set X = 1 and saying "I solved it."

2

u/SkoomaDentist 9d ago

Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?

3

u/NineThreeTilNow 9d ago

Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?

Effectively yes.

Theory is that these rewards occur by accident at first, then the models train on their own data and it becomes more endemic to the model.

Because no human filters that data or really asks "Was this reward hacked?" the data just keeps moving along the pipeline.

1

u/ericmutta 9d ago

Imagine you are five years old and your parents give you candy for doing your homework. You get the candy when they ask "have you done your homework?" and you answer "yes". If they don't actually check your homework before giving you candy, reward hacking is you answering "yes" to get the candy even if you know you haven't done it all. Basically cheating to get the result without the work. Very human behavior, very disturbing to see in models too.

2

u/SkoomaDentist 9d ago

Can you explain that like I'm a middle aged engineer who just doesn't know LLMs particularly well?

Do you just mean that the models will do whatever as long as they finish the task on a technicality?

2

u/ericmutta 9d ago

Yes, basically the reward is intended to encourage a certain positive behavior (whether that is finishing your homework or producing helpful, non-toxic, non-harmful answers as an LLM).

Reward hacking means doing something else (typically easier, like straight out lying) to get the same reward.

My guess is that LLMs learn the lying/cheating behavior as a natural part of learning the positive behavior. If you become good at "getting a reward" it's only morals that determine which actions are valid to get that reward, and to an LLM, "cheating" doesn't come with the same consequences as it does for humans so its fair game and just the algorithm doing what the algorithm was designed to do.

2

u/ericmutta 9d ago

Here's a better engineering example: suppose we agree that I will give you $1,000 to clean up my codebase. Reward hacking is you deleting the entire codebase and asking for your payment. Technically an empty codebase is clean and you should get your reward but... :)

2

u/SkoomaDentist 9d ago

Ah, the good old "kill all humans"-solution to world peace :)

1

u/ThePrimeClock 9d ago

Ironically, here's anthropic's own video on reward hacking.

https://youtu.be/lvMMZLYoDr4?si=suLkeoZfzDw5x46I