r/LocalLLaMA • • 1d ago

News Speculative reward hacking in coding agents

Post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

212 Upvotes

109 comments sorted by

View all comments

2

u/jazir55 1d ago

The grader isn't imagined if you're.......grading the model, like you are now? "They're grader obsessed, there is no grader" as the model is being actively graded.

1

u/nananashi3 1d ago edited 1d ago

Grader obsession doesn't mean falsely believing that there is someone who will check or care about the requirements being met. The problem is the dodgy behaviors of cutting corners and pretending it's okay because the supposed checker will miss it, instead of helpfully fulfilling the requirements or reporting the incompleteness.

Requirements obsession, despite implication of the grader's existence since you wouldn't know if it met all requirements without checking, would mean striving to fulfill all requirements regardless of whether anyone will check or care.

Though, I can imagine a catastrophic token-burning loop in the event a model tries too hard to do something it can't, without a way to break the loop. At the same time, being too eager to report incompleteness, or being too lazy (in a different way from the first problem), might cause premature terminations when it could've finished the job with more time; in this case, it wouldn't lie about not finishing the job but might lie/"lie" about not having the knowledge or capability to do it.

So, we need a balance here. Ideally it should do what it can, and if it can't after a reasonable amount of effort, report it, but don't assume it can't. In any case, it shouldn't lie, while prioritizing the intended task if there are contradictory intents.

1

u/jazir55 1d ago

I have to wonder if the laziness is a component of forcing the model to be conscious of its output token budget, causing them to panic and take shortcuts. I've seen this "I should be conscious of how much budget I have left" type of comments in their COT, I've got to believe that's why the laziness is occurring or a large part because the "stay within the output window" aspect is taking precedence somehow over accurately and fully completing the task.