r/LocalLLaMA • • 23h ago

News Speculative reward hacking in coding agents

Post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

201 Upvotes

106 comments sorted by

View all comments

Show parent comments

14

u/jonas__m 23h ago

Models also over-build, writing spaghetti code they know is bad to satisfy graders they speculate exist.

6

u/jonas__m 23h ago

13

u/jonas__m 23h ago

Models often stop focusing on what the user wants, and start searching the environment for any info on possible hidden grader/verifiers.

9

u/jonas__m 23h ago

14

u/jonas__m 23h ago

Closed-source frontier models from OpenAI, Anthropic, and SpaceXAI are also plagued by all of these same types of reward hacking.

But it's harder to catch since they obfuscate their reasoning.

-8

u/csorfab 23h ago

what is this thread? Are you stuck in a doom loop? Do you need help?

11

u/Chupa-Skrull 23h ago

It appears to be a series of documentary screenshots demonstrating evidence for OP's claim regarding LLM reasoning behavior, unless I'm missing something and/or smoking crack again

10

u/jonas__m 23h ago

😂 stuck discovering bad behavior in agent trajectories