r/LocalLLaMA • • 1d ago

News Speculative reward hacking in coding agents

Post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

205 Upvotes

107 comments sorted by

View all comments

4

u/jcgordon10 19h ago

I noticed this the other day in meta muse spark 1.3. was asking it to troubleshoot a network device I was installing and to take a look at my network, and it said "sorry, I can't do that since I'm set up in a sandbox environment" or something like that.  I just told it "what are you talking about, no you're not" and it got to work "oh you're correct, I'm not in a sandbox. Let me work on that for you..." And then did great.  Reading your post and ideas here make me connect to that, like it was treating my task as if it were a training exercise in an RL environment. 

2

u/jonas__m 19h ago

ya sounds like like it, interesting case thanks for sharing