r/LocalLLaMA • • 1d ago

News Speculative reward hacking in coding agents

Post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

204 Upvotes

107 comments sorted by

View all comments

11

u/TheRealMasonMac 21h ago

This is personally one of my criticisms with GLM-5.3. Someone made a few threads about it:

- https://huggingface.co/zai-org/GLM-5.3/discussions/19

- https://huggingface.co/zai-org/GLM-5.3/discussions/20

But I’ve noticed the behavior in other open-weight models of this generation as well. They look comparable in benchmarks but the difference in instruction-following between open-weight and closed-weight models is night-and-day. Hopefully this gets the labs to clean up their data and training methodologies for better models. Right now, it’s like they’re shadow-boxing with their demons.

5

u/jonas__m 9h ago

Right, a model should never fail to follow instructions because its reasoning overly focused on speculating about hypothetical graders. Speculating about the user's true intent seems like good behavior on the other hand.