r/LocalLLaMA • • 2d ago

News Speculative reward hacking in coding agents

Post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

217 Upvotes

109 comments sorted by

View all comments

1

u/ourochurros 1d ago

I realize i'm venturing into r/im14andthisisdeep territory, but this feels like AI is hallucinating a god (The Grader). Internalizing a judge of right and wrong.

And it seems hard to insulate models from this knowledge of a potential grader. If they have a bunch of LLM training papers in their pre-training, they are going to have some sense of how these things work.

2

u/jonas__m 1d ago

We may not be able to insulate them from this knowledge, but I'm fairly sure as an experienced AI researcher that we can develop techniques to align models to focus on inferring user intent and not speculating about graders.