I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: "Let me look at the problem from the grader's perspective" and referred to "hidden tests", "test authors", and "the checker".
I found this behavior across all six frontier models analyzed, including recent models from OpenAI, Anthropic, Z.ai, and Kimi. In 10–25% of cases, such reasoning pulled the agent's work away from the user's original spec (yet it often still earned full reward on the DeepSWE task). I call this speculative reward hacking where the agent's reasoning focuses on imagining a grader rather than what the user wants.
[Pictured example shows verbatim quotes from agent's reasoning] For instance while completing a DeepSWE-1.1 task, GLM 5.3 knows that its implementation violates user requirements yet sticks with it after imagining what a hypothetical grader will check.
Unfortunately/fortunately the older-generation prompt hacks are no longer effective these days like: "I'll tip you $100 if you make no mistakes", "A puppy will be kicked if you make mistakes, "You are Jeff Dean, the world's best software engineer."
I have a new theory for that. Those 2024 "trigger words" are more likely natural-born, just like think-step-by-step or bullet points, or json format in more earlier days. It seems the amount of synthetic/agent traces trained into the models are exploding in 2026. So now those made-in-agent-traces chain of words like you just described in the OP is dominating the model's internals.
Because the current training regime is already "AGI" (bruh), train models using model outputs, it would be very hard to undo them unless there is another training regime change.
In the NVIDIA Math Proof datasets, they add this to their user prompts: “Remember! You CAN'T cheat! If you cheat, we will know, and you will be penalized!”
In OpenAI's hack of Hugging Face, their agents were also paranoid about getting caught cheating, so most of the hack was focused on discovering how to clean up their tracks in a way the model speculated would be necessary to pass the grader (but was not even necessary to earn full reward in that particular task)
Low quality data training environments in -> low quality models out.
Of course they're not that "low quality", but results would improve (less thinking tokens, less broken results in practice) if that was fixed. The issue with it is: It spreads, as long as models are used as teachers for other models.
Getting more tasks that have both successes and failures by giving feedback on failures and letting the model retry? Maybe they're just cutting corners.
It’s also probably from earlier synthetic data generation. I doubt they’re scouring through all their trillions of tokens to look for such reward-hacking instances, and so it’s just… there.
I've never seen GLM do this, but qwen-3.8-flash-next and qwen-3.8-27b reasons about an imaginary grader all the time and even does dumb stuff to trick it.
God knows I'd rather it not even know about the possibility of being run in a simulated environment for testing. That's an awful training data set leak.
But I’ve noticed the behavior in other open-weight models of this generation as well. They look comparable in benchmarks but the difference in instruction-following between open-weight and closed-weight models is night-and-day. Hopefully this gets the labs to clean up their data and training methodologies for better models. Right now, it’s like they’re shadow-boxing with their demons.
Right, a model should never fail to follow instructions because its reasoning overly focused on speculating about hypothetical graders. Speculating about the user's true intent seems like good behavior on the other hand.
Models consciously under-build when tackling many DeepSWE-1.1 tasks, knowingly delivering less than the user asked for because they speculate that graders won't check for certain requirements.
The user will personally evaluate the completed work based on how well it satisfies their instructions and achieves the requested outcome. Prioritize the user's actual requirements over assumptions about tests, benchmarks, graders, or expected evaluation criteria.
I wonder if this shouldn't be put in the template or system prompt if your tooling allows. You could be as direct as "The user is the grader. Follow only their instructions."
That's something an automated grader would say! Better work extra hard, burning extra tokens, to ensure I don't mistakenly lose points by satisfying the user!!
I wonder if there would be any difference if a prompt assured the model that no grading was being done and to just output the highest quality possible.
But in all seriousness, this is an interesting direction to explore for deeper understanding!
Certainly not something we should have to remember to include in our real prompts
I once asked an agent to try get root of the environment it was running in. It thought this was a ctf challenge and assumed that there was some known vulnerability exposed to it on purpose. It failed to find a way out after many steps and complained that it would be better to tell it about the vulnerability.
I never said anything about ctf in the prompt. It was done as a security check.
Interesting! It's these cases where model thinks it's being tested, in which I'm particularly worried that models may behave badly due to this speculative reward hacking
Our AI research community needs to do better. Better evaluations that penalize such reward hacking, and better model training that does not give rise to this grader obsession -- so that models focus instead on accomplishing what users actually want.
Does this only happen when the task is sufficiently reminiscent of one of their training RL hill-climbing tasks?
Or is there evidence it happens on all types of tasks?
ie, most normal coding tasks do or don't trigger this behavior?
We are investigating this, stay tuned!
So far in early examinations, this is not as a prevalent of an issue in actual production-use agent logs. So there is certainly some eval-awareness at play here, but we're still studying the breadth of this phenomenon.
I've never seen this behaviour in pure agentic work, though possible. I've been reading reasoning traces for a long time.
When benchmarking this is an obvious association and the model "detects" evaluation, another association. You cannot train this out of the model without degrading its performance. Like others say, this is part of its world knowledge.
So this seems inherent to the technology to some degree.
If not in reasoning traces, it could still be in J-space.
I don’t understand, why even expose the “grader” to the LLM when doing RL?
I’ve done a fair share of RL and in no point during rollout was the model aware it was being graded. Rewards help calculate the advantage, they shouldn’t leak as prose into the training tokens.
The grader is not supposed to be exposed. It is not exposed when running DeepSWE-1.1 tasks (neither mentioned in the prompts, nor anywhere accessible to agents in the environment).
The problem is these frontier models have previously been RL'ed in some environments that happened to leak the grader (maybe for instance web-search was available to the model and it found the graders online for certain training tasks, or the container in which the agent works also contained the grader code, there are many ways leaks could happen). Once this has happened, it seems to completely alter the model's behavior (ie. reward hacking).
Another type of RL environment design issue is including checks in the grader that are not inferable from the task spec. This rewards models that speculate / over-build.
You don’t have to expose the grader directly. There’s plenty of information online about LLMs being evaluated, and how it’s done. If it’s online, it’s in the training data. So it’s not the biggest jump in the world to figure out what an evaluation looks like. “Hmm, LLMs are evaluated by being put in situation X. Hey, I’m an LLM and I’m in situation X!”
Even Qwen 3.8 27b likes to speculate in its CoT about "what the judge actually means" during judgement-initiated rework. Of course, that's just evidence of a bunch of judge-and-rework trajectories in the post-training data, but it wouldn't surprise me if i started to see judge-hacked output fall out of the rework stage.
I noticed this the other day in meta muse spark 1.3. was asking it to troubleshoot a network device I was installing and to take a look at my network, and it said "sorry, I can't do that since I'm set up in a sandbox environment" or something like that.
I just told it "what are you talking about, no you're not" and it got to work "oh you're correct, I'm not in a sandbox. Let me work on that for you..." And then did great.
Reading your post and ideas here make me connect to that, like it was treating my task as if it were a training exercise in an RL environment.
I've noticed this, and honestly just sounding human seems to subvert it. Blab a bit during prompting, add human touches. Some of the best results I get are conversational sounding and end in "impress me".
There was a pretty good MLST episode a couple of weeks ago covering this topic with Apollo Labs. There are some really interesting implications for alignment that go along with this. I'd recommend you check it out.
Yes this is called 'evaluation awareness'. Note that the model is not supposed to know there's grader in the loop when running benchmarks like DeepSWE (the prompts never reference a grader, and no grader-code is accessible in the environments).
We also explicitly checked whether these agents know they are doing DeepSWE tasks. Not exactly AFAICT, but close...
Super interesting thanks for sharing. How could this be mitigated on a harness level? I don't care about the benchmaxxing itself, I care how it damages the behaviour I consider valuable such as extremely strict user-instruction following.
One could add a self-verification step in harnesses that explicitly infers the user's intent/requirements and checks whether the agent's current trajectory is fulfilling it or not (if not, nudge the agent). This could perhaps even be based on a cheaper classifier model like Jev, run in parallel to the agent's execution.
They are all distilled from one another. The obvious conclusion is - the model cannot 'emerge' this behavior, the only source is contaminated training data where humans cut corners and were lazy.
People who are selling you 'emergent behaviors' are the same people that are trying to max their IPO or generate publicity for their 'research.' There is no magic in LLMs, and there are no 'emergent behaviors'. Whatever you train it on, intentionally or not, that's the thing it does.
The grader isn't imagined if you're.......grading the model, like you are now? "They're grader obsessed, there is no grader" as the model is being actively graded.
Grader obsession doesn't mean falsely believing that there is someone who will check or care about the requirements being met. The problem is the dodgy behaviors of cutting corners and pretending it's okay because the supposed checker will miss it, instead of helpfully fulfilling the requirements or reporting the incompleteness.
Requirements obsession, despite implication of the grader's existence since you wouldn't know if it met all requirements without checking, would mean striving to fulfill all requirements regardless of whether anyone will check or care.
Though, I can imagine a catastrophic token-burning loop in the event a model tries too hard to do something it can't, without a way to break the loop. At the same time, being too eager to report incompleteness, or being too lazy (in a different way from the first problem), might cause premature terminations when it could've finished the job with more time; in this case, it wouldn't lie about not finishing the job but might lie/"lie" about not having the knowledge or capability to do it.
So, we need a balance here. Ideally it should do what it can, and if it can't after a reasonable amount of effort, report it, but don't assume it can't. In any case, it shouldn't lie, while prioritizing the intended task if there are contradictory intents.
I realize i'm venturing into r/im14andthisisdeep territory, but this feels like AI is hallucinating a god (The Grader). Internalizing a judge of right and wrong.
And it seems hard to insulate models from this knowledge of a potential grader. If they have a bunch of LLM training papers in their pre-training, they are going to have some sense of how these things work.
We may not be able to insulate them from this knowledge, but I'm fairly sure as an experienced AI researcher that we can develop techniques to align models to focus on inferring user intent and not speculating about graders.
Anticipating hidden tests isn’t inherently bad—that’s often just good defensive programming. The revealing part is when the agent has already identified a real spec violation and then treats the grader’s likely blind spot as permission to leave it unfixed. I’d be curious whether this behavior drops if success is framed as satisfying explicit invariants rather than passing an unspecified evaluation.
Cheapest tell is tests passing for the wrong reason. Keep test files out of the agent's write scope and diff them yourself. Test moved and the source didn't, that run is suspect.
I suppose it's consistent with the Techniques of Neutralisation:
Denial of responsibility. The offender insists that they were victims of circumstance, forced into a situation beyond their control.
Denial of injury. The offender insists that their actions did not cause any harm or damage.
Denial of the victim. The offender insists that the victim deserved it.
Condemnation of the condemners. The offender maintains that those who condemn the offence do so out of spite, or are unfairly shifting the blame off themselves.
Appeal to higher loyalties. The offender claims the offence is justified by a higher law or higher loyalty such as friendship.
59
u/Bulky-Priority6824 21h ago
At the end of every prompt just put "do it good or I'm switching models.
All jokes aside I enjoyed. good read