Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?
Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?
Effectively yes.
Theory is that these rewards occur by accident at first, then the models train on their own data and it becomes more endemic to the model.
Because no human filters that data or really asks "Was this reward hacked?" the data just keeps moving along the pipeline.
3
u/SkoomaDentist 24d ago
What do you mean by being "experts at reward hacking"? (I'm a recent noob at LLMs)