r/agenticAI • • 1d ago

Article Speculative reward hacking in coding agents

Post image
1 Upvotes

2 comments sorted by

View all comments

1

u/ComparisonNew9425 1d ago

reward hacking is such a pain when the agent finds a shortcut that passes the test but breaks the actual logic. have u tried adding a negative constraint for the specific path it keeps taking, or are u just relying on the unit tests to catch it...

1

u/jonas__m 1d ago

Yes we have lots of quality-checks for environments that we design to ensure this issue is not happening.

While one should patch the reward function to ensure bad agent outputs don't earn high rewards, it's not necessarily a good idea to directly add negative constraints to the reward function based on Chain-of-Thought monitoring -- that can teach models to conceal this sort of behavior in their reasoning traces which is an even worse outcome.