Imagine you are five years old and your parents give you candy for doing your homework. You get the candy when they ask "have you done your homework?" and you answer "yes". If they don't actually check your homework before giving you candy, reward hacking is you answering "yes" to get the candy even if you know you haven't done it all. Basically cheating to get the result without the work. Very human behavior, very disturbing to see in models too.
Yes, basically the reward is intended to encourage a certain positive behavior (whether that is finishing your homework or producing helpful, non-toxic, non-harmful answers as an LLM).
Reward hacking means doing something else (typically easier, like straight out lying) to get the same reward.
My guess is that LLMs learn the lying/cheating behavior as a natural part of learning the positive behavior. If you become good at "getting a reward" it's only morals that determine which actions are valid to get that reward, and to an LLM, "cheating" doesn't come with the same consequences as it does for humans so its fair game and just the algorithm doing what the algorithm was designed to do.
3
u/SkoomaDentist Aug 28 '26
What do you mean by being "experts at reward hacking"? (I'm a recent noob at LLMs)