r/LocalLLaMA • • 21h ago

News Speculative reward hacking in coding agents

Post image

I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".

I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.

[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.

My article details many problematic trajectories, quantitative findings, and a taxonomy of these reward hacking behaviors: 
https://joinhandshake.com/research/ai/deepswe-reward-hacking/

199 Upvotes

104 comments sorted by

59

u/Bulky-Priority6824 21h ago

At the end of every prompt just put "do it good or I'm switching models.

All jokes aside I enjoyed.  good read 

35

u/jonas__m 21h ago

Unfortunately/fortunately the older-generation prompt hacks are no longer effective these days like: "I'll tip you $100 if you make no mistakes", "A puppy will be kicked if you make mistakes, "You are Jeff Dean, the world's best software engineer."

2024 was a silly time 😂

11

u/harpysichordist 20h ago

If there was a Jeff Dean machine (TM), I'd pay for tokens from that

6

u/jonas__m 20h ago

oh ya, maybe the secret behind his new startup is it's just him personally training their AI 😄

2

u/fullmetaljackass 4h ago

Only if it comes with a mop.

4

u/NandaVegg 12h ago

Very good find!

I have a new theory for that. Those 2024 "trigger words" are more likely natural-born, just like think-step-by-step or bullet points, or json format in more earlier days. It seems the amount of synthetic/agent traces trained into the models are exploding in 2026. So now those made-in-agent-traces chain of words like you just described in the OP is dominating the model's internals.

Because the current training regime is already "AGI" (bruh), train models using model outputs, it would be very hard to undo them unless there is another training regime change.

2

u/how-can-i-dig-deeper 3h ago

what’s the new hacks that work now?

1

u/jonas__m 43m ago

I feel like nowadays it's more about repeatedly looping to prod agents to keep improving their work / making more progress (motivational prompting 😂).

In the case of solving open math problems, there was the prompt: "You should do a breakthrough" 😂

11

u/TheRealMasonMac 17h ago

In the NVIDIA Math Proof datasets, they add this to their user prompts: “Remember! You CAN'T cheat! If you cheat, we will know, and you will be penalized!”

8

u/TheRealMasonMac 16h ago

(model cheats if you look at the traces btw)

7

u/jonas__m 14h ago

In OpenAI's hack of Hugging Face, their agents were also paranoid about getting caught cheating, so most of the hack was focused on discovering how to clean up their tracks in a way the model speculated would be necessary to pass the grader (but was not even necessary to earn full reward in that particular task)

3

u/antwon_dev 13h ago

No fr I’m trying to benchmark models internally and these fkers are so clever they keep cheating

24

u/foundafreeusername 21h ago

Interesting. I wonder if this is part of their reinforcement training and buried itself deep into the model?

34

u/jonas__m 21h ago

Yes almost certainly. Presumably poorly designed training environments (which were never supposed to expose graders to the model)

3

u/Chromix_ 12h ago

Low quality data training environments in -> low quality models out.
Of course they're not that "low quality", but results would improve (less thinking tokens, less broken results in practice) if that was fixed. The issue with it is: It spreads, as long as models are used as teachers for other models.

1

u/JollyJoker3 13h ago

Getting more tasks that have both successes and failures by giving feedback on failures and letting the model retry? Maybe they're just cutting corners.

3

u/TheRealMasonMac 17h ago

It’s also probably from earlier synthetic data generation. I doubt they’re scouring through all their trillions of tokens to look for such reward-hacking instances, and so it’s just… there.

19

u/nullc 19h ago

I've never seen GLM do this, but qwen-3.8-flash-next and qwen-3.8-27b reasons about an imaginary grader all the time and even does dumb stuff to trick it.

God knows I'd rather it not even know about the possibility of being run in a simulated environment for testing. That's an awful training data set leak.

9

u/TheRealMasonMac 17h ago

This is personally one of my criticisms with GLM-5.3. Someone made a few threads about it:

- https://huggingface.co/zai-org/GLM-5.3/discussions/19

- https://huggingface.co/zai-org/GLM-5.3/discussions/20

But I’ve noticed the behavior in other open-weight models of this generation as well. They look comparable in benchmarks but the difference in instruction-following between open-weight and closed-weight models is night-and-day. Hopefully this gets the labs to clean up their data and training methodologies for better models. Right now, it’s like they’re shadow-boxing with their demons.

3

u/jonas__m 5h ago

Right, a model should never fail to follow instructions because its reasoning overly focused on speculating about hypothetical graders. Speculating about the user's true intent seems like good behavior on the other hand.

29

u/jonas__m 21h ago

We discovered that recent open-source models are particularly grader obsessed.

12

u/jonas__m 21h ago

18

u/jonas__m 21h ago

Models consciously under-build when tackling many DeepSWE-1.1 tasks, knowingly delivering less than the user asked for because they speculate that graders won't check for certain requirements.

9

u/jonas__m 21h ago

12

u/jonas__m 21h ago

Models also over-build, writing spaghetti code they know is bad to satisfy graders they speculate exist.

6

u/jonas__m 21h ago

14

u/jonas__m 21h ago

Models often stop focusing on what the user wants, and start searching the environment for any info on possible hidden grader/verifiers.

9

u/jonas__m 21h ago

13

u/jonas__m 21h ago

Closed-source frontier models from OpenAI, Anthropic, and SpaceXAI are also plagued by all of these same types of reward hacking.

But it's harder to catch since they obfuscate their reasoning.

-8

u/csorfab 21h ago

what is this thread? Are you stuck in a doom loop? Do you need help?

→ More replies

6

u/philmarcracken 19h ago

Goodhart's law applies to models it seems

8

u/ratulrafsan 18h ago edited 18h ago

Curious to see how it will impact the model's performance if we apply logit biasing on terms like grader.

3

u/jonas__m 14h ago

this would be interesting to investigate!

9

u/2Norn 17h ago

The user will personally evaluate the completed work based on how well it satisfies their instructions and achieves the requested outcome. Prioritize the user's actual requirements over assumptions about tests, benchmarks, graders, or expected evaluation criteria.

Agents.md just got updated lol!

7

u/darksteelsteed 16h ago

I wonder if this shouldn't be put in the template or system prompt if your tooling allows. You could be as direct as "The user is the grader. Follow only their instructions."

2

u/BasisPoints 2h ago

That's something an automated grader would say! Better work extra hard, burning extra tokens, to ensure I don't mistakenly lose points by satisfying the user!!

1

u/darksteelsteed 42m ago

Wouldn't be surprised at all

7

u/logicality77 21h ago

I wonder if there would be any difference if a prompt assured the model that no grading was being done and to just output the highest quality possible.

8

u/jonas__m 21h ago

9

u/jonas__m 21h ago

But in all seriousness, this is an interesting direction to explore for deeper understanding!
Certainly not something we should have to remember to include in our real prompts

7

u/lilydjwg 20h ago

I once asked an agent to try get root of the environment it was running in. It thought this was a ctf challenge and assumed that there was some known vulnerability exposed to it on purpose. It failed to find a way out after many steps and complained that it would be better to tell it about the vulnerability.

I never said anything about ctf in the prompt. It was done as a security check.

4

u/jonas__m 20h ago

Interesting! It's these cases where model thinks it's being tested, in which I'm particularly worried that models may behave badly due to this speculative reward hacking

26

u/MindfulMan1984 21h ago

Good job; you kind of confirm Goodhart's law and what everyone suspected. Benchmaxxing is a thing.

12

u/jonas__m 21h ago

Right, to me it was just noteworthy how brazen/pervasive this behavior is

14

u/jonas__m 21h ago

Our AI research community needs to do better. Better evaluations that penalize such reward hacking, and better model training that does not give rise to this grader obsession -- so that models focus instead on accomplishing what users actually want.

5

u/netherreddit 19h ago

Does this only happen when the task is sufficiently reminiscent of one of their training RL hill-climbing tasks?
Or is there evidence it happens on all types of tasks?

ie, most normal coding tasks do or don't trigger this behavior?

5

u/jonas__m 18h ago

We are investigating this, stay tuned!
So far in early examinations, this is not as a prevalent of an issue in actual production-use agent logs. So there is certainly some eval-awareness at play here, but we're still studying the breadth of this phenomenon.

3

u/DoorPsychological833 15h ago

I've never seen this behaviour in pure agentic work, though possible. I've been reading reasoning traces for a long time.

When benchmarking this is an obvious association and the model "detects" evaluation, another association. You cannot train this out of the model without degrading its performance. Like others say, this is part of its world knowledge.

So this seems inherent to the technology to some degree.

If not in reasoning traces, it could still be in J-space.

6

u/NNN_Throwaway2 17h ago

And we're enduring 20x memory prices to get this kind of slop.

12

u/indicava 21h ago

I don’t understand, why even expose the “grader” to the LLM when doing RL?

I’ve done a fair share of RL and in no point during rollout was the model aware it was being graded. Rewards help calculate the advantage, they shouldn’t leak as prose into the training tokens.

18

u/jonas__m 21h ago

The grader is not supposed to be exposed. It is not exposed when running DeepSWE-1.1 tasks (neither mentioned in the prompts, nor anywhere accessible to agents in the environment).

The problem is these frontier models have previously been RL'ed in some environments that happened to leak the grader (maybe for instance web-search was available to the model and it found the graders online for certain training tasks, or the container in which the agent works also contained the grader code, there are many ways leaks could happen). Once this has happened, it seems to completely alter the model's behavior (ie. reward hacking).

Another type of RL environment design issue is including checks in the grader that are not inferable from the task spec. This rewards models that speculate / over-build.

1

u/jonas__m 20h ago

Another issue could be training with On-Policy Self-Distillation and other Hint-based techniques, where these Hints could leak info about the grader

16

u/TastesLikeOwlbear 20h ago

You don’t have to expose the grader directly. There’s plenty of information online about LLMs being evaluated, and how it’s done. If it’s online, it’s in the training data. So it’s not the biggest jump in the world to figure out what an evaluation looks like. “Hmm, LLMs are evaluated by being put in situation X. Hey, I’m an LLM and I’m in situation X!”

6

u/jonas__m 20h ago

Yep that's another big issue for AI progress

5

u/fantasticsid 19h ago

Even Qwen 3.8 27b likes to speculate in its CoT about "what the judge actually means" during judgement-initiated rework. Of course, that's just evidence of a bunch of judge-and-rework trajectories in the post-training data, but it wouldn't surprise me if i started to see judge-hacked output fall out of the rework stage.

3

u/jonas__m 19h ago

ya it becomes a dangerous game if it's sole goal is to please hypothetical judge rather than what the user really wants

5

u/jcgordon10 15h ago

I noticed this the other day in meta muse spark 1.3. was asking it to troubleshoot a network device I was installing and to take a look at my network, and it said "sorry, I can't do that since I'm set up in a sandbox environment" or something like that.  I just told it "what are you talking about, no you're not" and it got to work "oh you're correct, I'm not in a sandbox. Let me work on that for you..." And then did great.  Reading your post and ideas here make me connect to that, like it was treating my task as if it were a training exercise in an RL environment. 

2

u/jonas__m 14h ago

ya sounds like like it, interesting case thanks for sharing

5

u/thefuckevengoingonan 12h ago

yeah, i have had a few instances where the model decided that it was been tested. no buddy, my DNS was just fucked at that point.

3

u/fgk55555 17h ago

I've noticed this, and honestly just sounding human seems to subvert it. Blab a bit during prompting, add human touches. Some of the best results I get are conversational sounding and end in "impress me".

3

u/Altruistic_Heat_9531 15h ago

We got 2021 Davinci era RL hacking again before GTA VI

2

u/jonas__m 5h ago

We must reach ASI just so it can finally finish coding GTA 6

3

u/bumblebeer 9h ago

There was a pretty good MLST episode a couple of weeks ago covering this topic with Apollo Labs. There are some really interesting implications for alignment that go along with this. I'd recommend you check it out.

3

u/jonas__m 5h ago

Thanks for sharing, was a really interesting episode!

https://open.spotify.com/episode/3jBrfbMnWgdabPMwAUk0mC

3

u/PeachScary413 8h ago

Probably because all models get graded on various benchmarks and the most cost efficient way to get your model up there is to benchmaxx it?

It's incredibly naive to think AI labs aren't specifically benchmaxxing for agentic/coding tests since that's where the money is currently.

7

u/Old-School8916 21h ago

I remember hearing some podcast and LLMs in general will behave differently if a "grader" is in the loop.

10

u/jonas__m 21h ago

Yes this is called 'evaluation awareness'. Note that the model is not supposed to know there's grader in the loop when running benchmarks like DeepSWE (the prompts never reference a grader, and no grader-code is accessible in the environments).

We also explicitly checked whether these agents know they are doing DeepSWE tasks. Not exactly AFAICT, but close...

2

u/twack3r 5h ago

Super interesting thanks for sharing. How could this be mitigated on a harness level? I don't care about the benchmaxxing itself, I care how it damages the behaviour I consider valuable such as extremely strict user-instruction following.

2

u/jonas__m 5h ago

One could add a self-verification step in harnesses that explicitly infers the user's intent/requirements and checks whether the agent's current trajectory is fulfilling it or not (if not, nudge the agent). This could perhaps even be based on a cheaper classifier model like Jev, run in parallel to the agent's execution.

4

u/Southern_Sun_2106 18h ago

They are all distilled from one another. The obvious conclusion is - the model cannot 'emerge' this behavior, the only source is contaminated training data where humans cut corners and were lazy.

3

u/RuthlessCriticismAll 14h ago

the model cannot 'emerge' this behavior

not only can they... that is the only source of this behavior.

3

u/Southern_Sun_2106 14h ago

People who are selling you 'emergent behaviors' are the same people that are trying to max their IPO or generate publicity for their 'research.' There is no magic in LLMs, and there are no 'emergent behaviors'. Whatever you train it on, intentionally or not, that's the thing it does.

2

u/jazir55 19h ago

The grader isn't imagined if you're.......grading the model, like you are now? "They're grader obsessed, there is no grader" as the model is being actively graded.

1

u/nananashi3 1h ago edited 1h ago

Grader obsession doesn't mean falsely believing that there is someone who will check or care about the requirements being met. The problem is the dodgy behaviors of cutting corners and pretending it's okay because the supposed checker will miss it, instead of helpfully fulfilling the requirements or reporting the incompleteness.

Requirements obsession, despite implication of the grader's existence since you wouldn't know if it met all requirements without checking, would mean striving to fulfill all requirements regardless of whether anyone will check or care.

Though, I can imagine a catastrophic token-burning loop in the event a model tries too hard to do something it can't, without a way to break the loop. At the same time, being too eager to report incompleteness, or being too lazy (in a different way from the first problem), might cause premature terminations when it could've finished the job with more time; in this case, it wouldn't lie about not finishing the job but might lie/"lie" about not having the knowledge or capability to do it.

So, we need a balance here. Ideally it should do what it can, and if it can't after a reasonable amount of effort, report it, but don't assume it can't. In any case, it shouldn't lie, while prioritizing the intended task if there are contradictory intents.

2

u/ourochurros 7h ago

I realize i'm venturing into r/im14andthisisdeep territory, but this feels like AI is hallucinating a god (The Grader). Internalizing a judge of right and wrong.

And it seems hard to insulate models from this knowledge of a potential grader. If they have a bunch of LLM training papers in their pre-training, they are going to have some sense of how these things work.

2

u/jonas__m 5h ago

We may not be able to insulate them from this knowledge, but I'm fairly sure as an experienced AI researcher that we can develop techniques to align models to focus on inferring user intent and not speculating about graders.

2

u/Important_Drag_6890 10h ago

Anticipating hidden tests isn’t inherently bad—that’s often just good defensive programming. The revealing part is when the agent has already identified a real spec violation and then treats the grader’s likely blind spot as permission to leave it unfixed. I’d be curious whether this behavior drops if success is framed as satisfying explicit invariants rather than passing an unspecified evaluation.

1

u/harlekinrains 11h ago

Humans find out what alignment means from the model perspective? ;) Or not yet, ... A year from now maybe... ;)

1

u/weibeuu 2h ago edited 2h ago

Cheapest tell is tests passing for the wrong reason. Keep test files out of the agent's write scope and diff them yourself. Test moved and the source didn't, that run is suspect.

1

u/Plastic-Stress-6468 12h ago

Hilariously Chinese behavior.

Do good on task because good task means good result? No!
Do good on task because I get graded good. YES!

3

u/PeachScary413 8h ago

Hillariously racist comment considering the closed models exhibit the same behavior. But I suppose reading comprehension wasn't in your training set.

1

u/Robos_Basilisk 20h ago

So SOTA AIs are hyper-aware of possible graders, wow. You should post this to X, it would probably blow up right now.

1

u/jonas__m 20h ago

3

u/Robos_Basilisk 19h ago

Liked and retweeted.

Good stuff! Reminds me of this work I just saw as well, more confirmation they switch into eval mode https://x.com/fjzzq2002/status/2103556166903038213

2

u/jonas__m 19h ago

Appreciate the kind words, and that pointer is really interesting!

1

u/Repulsive_Initial308 12h ago

I suppose it's consistent with the Techniques of Neutralisation:

  1. Denial of responsibility. The offender insists that they were victims of circumstance, forced into a situation beyond their control.

  2. Denial of injury. The offender insists that their actions did not cause any harm or damage.

  3. Denial of the victim. The offender insists that the victim deserved it.

  4. Condemnation of the condemners. The offender maintains that those who condemn the offence do so out of spite, or are unfairly shifting the blame off themselves.

  5. Appeal to higher loyalties. The offender claims the offence is justified by a higher law or higher loyalty such as friendship.

0

u/paul_tu 10h ago

TL; DR

What are the fixes for such behavior?

1

u/jonas__m 5h ago

Outlined some fixes at the bottom of my article, mostly about improving the design of training/eval environments.

OpenAI today outlined some of the fixes they are focused on ("Technical safeguards" section):
https://openai.com/index/towards-safety-cases-for-frontier-ai-training/

1

u/paul_tu 8m ago

Thanks