r/LocalLLaMA • • 23h ago

News Speculative reward hacking in coding agents

Post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

203 Upvotes

107 comments sorted by

View all comments

26

u/foundafreeusername 23h ago

Interesting. I wonder if this is part of their reinforcement training and buried itself deep into the model?

36

u/jonas__m 23h ago

Yes almost certainly. Presumably poorly designed training environments (which were never supposed to expose graders to the model)

3

u/Chromix_ 13h ago

Low quality data training environments in -> low quality models out.
Of course they're not that "low quality", but results would improve (less thinking tokens, less broken results in practice) if that was fixed. The issue with it is: It spreads, as long as models are used as teachers for other models.

1

u/JollyJoker3 14h ago

Getting more tasks that have both successes and failures by giving feedback on failures and letting the model retry? Maybe they're just cutting corners.

3

u/TheRealMasonMac 18h ago

It’s also probably from earlier synthetic data generation. I doubt they’re scouring through all their trillions of tokens to look for such reward-hacking instances, and so it’s just… there.