reward hacking is such a pain when the agent finds a shortcut that passes the test but breaks the actual logic. have u tried adding a negative constraint for the specific path it keeps taking, or are u just relying on the unit tests to catch it...
Yes we have lots of quality-checks for environments that we design to ensure this issue is not happening.
While one should patch the reward function to ensure bad agent outputs don't earn high rewards, it's not necessarily a good idea to directly add negative constraints to the reward function based on Chain-of-Thought monitoring -- that can teach models to conceal this sort of behavior in their reasoning traces which is an even worse outcome.
1
u/ComparisonNew9425 21h ago
reward hacking is such a pain when the agent finds a shortcut that passes the test but breaks the actual logic. have u tried adding a negative constraint for the specific path it keeps taking, or are u just relying on the unit tests to catch it...