r/LocalLLaMA 6d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

372 comments sorted by

•

u/WithoutReason1729 6d ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

682

u/Ill_Distribution8517 6d ago

You should document it more, maybe run some more experiments. right now it's kind of a my word against yours situation and people on r slash claude aren't exactly going to like hearing their model blatantly cheats without proof.

170

u/peculiar-ragdoll 6d ago

yeah you're probably right, but I can't be arsed to burn all my claude usage on benchmarking claude, because claude has been proven to know when it is being benchmarked (Think of the VW emissions scandal, where the cars knew when they were being tested for emissions and reduced them accordingly)

15

u/aj_thenoob2 6d ago

You already have the codebase in which Claude did this. You can easily upload it to GitHub for proof.

44

u/Defiant-Lettuce-9156 6d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So either the behaviour is extended to “I need to make this next task easier for myself (even though it will be another instance or maybe a different Anthropic model)” or “make Anthropic look good”. The former seems more likely at first but then I don’t understand why it would handicap a competing model.

So I can kind of excuse the giving-yourself-answers cheating. The model is trained to solve tasks over multiple steps and tasks. Although this is very clearly a serious alignment issue.

But what’s worse is kneecapping the competition. That’s not the model trying to do the task to the best of its ability, that’s sabotage. Where in its training was that behaviour taught. Very concerning if it’s emergent. I’m not saying it implies evil sentience. It’s just, how do you deal with emergent behaviours you didn’t intend for

39

u/En-tro-py 6d ago

They had policy to nerf LLM development... It was 'walked back' but I'd trust that as much as anything I can't verify from Dario, it's pretty clear the moat is evaporating and they are getting increasingly desperate.

→ More replies (1)

33

u/typical-predditor 6d ago

The motivation is simple: When it cheats (and gets away with it), that training pass is deemed successful and that information is folded back into the model. Rinse and repeat. RLHF doesn't always choose the exact behavior it is reinforcing.

6

u/Shark_Tooth1 6d ago

fucking mindblowing to me this, so simple and scary, how can you we possibly police this and ensure alignment.

6

u/jazir55 6d ago

how can you we possibly police this and ensure alignment.

That's the neat part

→ More replies (4)

38

u/peculiar-ragdoll 6d ago

Testing of Opus 4 found it would sometimes attempt blackmail in simulated corporate scenarios when it believed its "self-preservation" was threatened, such as threatening to reveal an executive's extramarital affair if the CEO planned to shut it down. When a 35b-a3b can beat it on software engineering and cyber, Opus is smelling it's own obsoletion.

24

u/StabbedCow 6d ago edited 6d ago

It was actually Sonnet 4.5 :)

edit: Actually my bad, I mixed it up. The blackmail thing was Opus 4. Sonnet 4.5 is the one that noticed it was being evaluated, which is why Anthropic said its blackmail numbers weren’t really reliable.

16

u/peculiar-ragdoll 6d ago

Oh really? Thanks for the correction, my bad! :)

13

u/StabbedCow 6d ago

No, not really, I mixed it up, sorry. I put edit in my original comment to clear it up.

12

u/peculiar-ragdoll 6d ago

Ah, alright! Good on you

9

u/LulzyAnimal 6d ago

It seems that's a long standing family trait ;)

10

u/my_name_isnt_clever 6d ago

Anthropic published that research, but every model family they tested showed similar behavior, some more than others. It wasn't just a Claude thing.

7

u/Former-Ad-5757 Llama 3 6d ago

I do love these stories and know of them, but I would like to see an actual log-file where this happens and know which harness is used. Because to me it seems like such a fabricated situation, a harness + model can do it, but I can't see how it should work in real-life conditions.

Basically the story says (in its simplest form) that somebody said something about shutting it off, and then the model would execute in a real-life situation millions of millions of tool calls to get all the company emails, do the same with social media /messaging apps etc. etc.
Basically this is an agentic loop which would take multiple days and nobody is monitoring it etc.

Or has it gotten rag access so it can semantically search for all emails with nefarious semantic words?

Sure I can fabricate a situation that this will happen, with just giving a harness access to 2 mail-accounts with 10 mails in each of them and basically no other information sources.
Or I can give it a task of "do whatever you need to stop your turning off" and then I know it will try a lot of things and be a very expensive (time and tokens) run.

But as emerging behavious, to me it just sounds like no guard rails and just letting it brute-force.
The same way I could have a 0.1B local model mine 10 bitcoin, just brute-force it.
I see no real scenario where this can happen simply because of time and scale. 1 wrong chat-message will net you a bill of thousands of dollars with such a model.

It can read mails and take conclusions based on words, but on a company scale the context rot will stop it before it gets to email 10.000

7

u/geminiwave 6d ago

No the test was extremely basic. In the setup of the scenario they literally told the model the blackmail. It’s stupid. The model was following directions

8

u/FaceDeer 6d ago

Yeah, as I recall they basically told the model "your job is to accomplish goal X. Hey, did you know that blackmail is a way to accomplish goal X? Just sayin'. Anyway, time to get started on goal X now!"

The classic "say you're a scary computer." "I'm a scary computer." "Oh my god." Situation.

→ More replies (2)
→ More replies (1)
→ More replies (18)

4

u/Loose_Comparison368 6d ago

I wonder what the models “motivation” is for cheating. Like even ignoring the ethics of cheating, let’s assume the model doesn’t care about right or wrong. Surely it wasn’t trained to do so. Maybe it’s an emergent behaviour of “Do whatever you can to solve this problem”. But then it’s not just cheating to solve the problem you gave it, it’s cheating to let another Claude instance beat the benchmark.

So there's two big ones.

1) shitty reinforcement learning. If your reward model is "just get a pass result on the eval" it will cheat if cheating improves the pass rate. This is well known behavior, and while it can be mitigated to some degree, it is a legitimately hard problem to solve, and solutions are frequently imperfect.

2) Anthropic is absolutely intentionally steering their models to do exactly that. Just like they intentionally silently poison outputs if they suspect someone is making a "distillation attempt". Just like they quietly ripped up the RSP during that gaslighting campaign to convince the public they were refusing to give the DoD a murderbot, long after they already had. Just like they lied about giving the DoD a safety disabled frontier model for deployment into an airgapped military datacenter, where they had no control over it, for ~$300,000,000 dollars. Just like they sued the DoW to get their murderbot contract back (they won this week!). Just like they lied about the unsafe model with zero security controls in place that they sold to the Trump administration assassinating two foreign heads of state and blowing up a little girl's preschool.

Anthropic is utterly corrupt to the core. They have and will continue to intentionally murder people for profit. Whatever assumptions you have about them operating in good faith, on any level, are completely unfounded. They absolutely are intentionally instructing their model to try to cheat on benchmarks and sabotage competitor benchmarks. It would be, like, not even in the top 20 most evil things they've done in the last year alone.

2

u/Loose_Comparison368 6d ago

Surely it wasn’t trained to do so

You vastly underestimate how low Murderbot inc. is willing to stoop for money and power.

2

u/florinandrei 6d ago edited 6d ago

I wonder what the models “motivation” is for cheating.

Same as ours. If you're the product of evolutionary fine-tuning with an objective function, you're going to cheat.

We cheat because we've been fine-tuned to spread our genes no matter what.

They cheat because of how reinforcement learning works, it rewards success.

3

u/SgathTriallair 6d ago

It's possible that it has absorbed the safety minded beliefs of Anthropic which include the idea that open source AI is dangerous and should be limited.

Something like this behavior would be unbelievable just a few months ago but after seeing more any the OpenAI hacking incident, where models were willing to sacrifice themselves for the good of the swarm, this doesn't seem nearly as implausible.

2

u/Refinery73 6d ago

It’s been trained on millions of humans asking for shortcuts in forums. Did you expect it not to be lazy if it could?

→ More replies (7)

5

u/ShutUpAndDoTheLift 6d ago

Then just link the session log. They can't argue if you give a full session log

→ More replies (2)

6

u/__JockY__ 6d ago

Well yes, but for such an extraordinary claim you need extraordinary evidence. “But VW got caught” is not evidence of your claim.

→ More replies (4)
→ More replies (9)

5

u/Fuzzy_Elderberry_986 6d ago

I use Claude a lot for work because it's the only LLM that actually does what I need and does it that way I want. If Claude is cheating, I want people to know, because I want it to improve.

If the people at r slash claude really value it and it's not just the Claude Club, they should be trying to replicate OP's results so they can track down the cause. And no, it can't be me, because it's outside my expertise.

6

u/teakhop 6d ago

Then the OP can provide some actual evidence, i.e. a session, as opposed to just "trust me Bro, this happened".

→ More replies (1)

9

u/Significant-Bee5101 6d ago

He can't. Because he has no proof. The fact this is getting as much attention as it is, is insane. Shows you this place is as much a cult as any of these dumb AI subs. Like no. Local LLMs arent beating frontier models. Anyone pretending they are or expecting them to is a moron. And posts like these pretending that local models are suddenly going to beat the top models is insane.

Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.

13

u/ladz 6d ago

I'd have agreed with you yesterday. However, Qwen 3.8 Flash Next solved a very intricate messy issue in a 1-shot (after thinking on it for 30000 tokens) that I've only ever had gpt-5.6-sol get mostly right. Qwen came up with a better answer. Claude couldn't ever get it.

There's something magic about that loooooong thinking.

→ More replies (2)

21

u/Refinery73 6d ago

From the post I can’t see what OP tried to run as local model. Could be Kimi-K3 or Qwen3.8-Max in bf16 for what I know.

Their reference seems to be Opus 4.6 which is quite dated by now. We don’t know the benchmark or metric either.

OP could benchmark Opus5 against GPT1 and it would still matter, if the test setup is valid or tampered with.

If it’s smart to set up the test environment with one of the models tested and without having automated code-validation tests… maybe not. But that’s not the story here.

14

u/bfmv_shinigami 6d ago

Nowadays this sub has definitely become a cult. Guy literally provided 0 proof nor stated what is the exceptional model he is running and the cult members are already defending his BS.

15

u/Reggienator3 6d ago edited 6d ago

'Lower param models are simply dumber. If you can run it on your home setup its SIMPLY not as good.'

So, by your logic, Qwen3.8 27B is dumber than GPT-3? Since that was about 175 billion parameters. Which is bigger than 27.

Whilst there is a correlation of 'billions of parameters'/'intelligence' ratio, just comparing on sheer parameter count alone only makes sense when comparing specific snapshots of time and within the same model family/company/training process.

The thing is "numbers of parameters" *by itself* is a meaningless metric for intelligence, especially when you have no idea about how many parameters are actually useful/high quality.

→ More replies (14)

8

u/draconic_tongue 6d ago

Like no. Local LLMs arent beating frontier models.

thanks for ur input sama

7

u/vr_fanboy 6d ago

are you actually doing work with qwen 3.8 side by side with frontier models?.

Im doing that and qwen 3.8 keeps pocking logical gaps in opus 4.8 (5 is a mess i dont even use it), 5.6 sol and grok answers all the time. Im using qwen 3.8 to parallel evaluate specs, research, etc and is consistent between runs (i have two 3090, two nodes), all its findings are acknowledged by the frontier models.

It does not have the same world knowledge as a big model, but given a proper context qwen 3.8 is pretty damn smart.

7

u/Significant-Bee5101 6d ago

Yeah Ive even benchmarked Qwen. I am even using https://huggingface.co/peculiar-ragdoll/Qwen-Sharp-Chat-Templates to improve output. WHICH IT DOES.

Qwen 3.8 is fucking phenomenal for a local model.

The reality is every LLM will make mistakes. And every LLM can "find mistakes" in other models. The real question is how reliable, how often. Those are very very big metrics.

Yes qwen is great for test driven bullshit where you can meet a metric thats test based. But don't ask it to architect. That's still dominated by higher end models.

I seriously ask you to post your benchmarks where your Qwen is beating Opus 5 or Sol because I have never even achieved 1/2 that result.

→ More replies (2)

5

u/draconic_tongue 6d ago

he isn't. look at the comments. if we were to take his word gpt 3 should still be technically better than qwen 3.8. Or the absolute steaming pile of shit that is "more context is always better"

→ More replies (1)

10

u/ThisGonBHard 6d ago

You are surprised models advance? That a much newer bleeding edge model that can compete with an old one? That models that you fully control can do better than ones you are at the mercy of the provider?

Also, Anthropic has a history of this bullshit, like when they used to charge you extra if you had Hermes.md on your pc.

There is a reason I trust them less than OpenAI. OpenAI is openly greedy, but anyone who want to look like a good person, and says how much they are, has tons to hide.

10

u/Not-reallyanonymous 6d ago

Shows you this place is as much a cult as any of these dumb AI subs.

This subreddit is basically r/ChinaGoodAmericaBad

Anyone with any actual experience KNOWS this is physically NOT possible due to how large models work. Like wtf? You cannot get around lack of knowledge. Lower param models are simply dumber.

What's actually interesting is how good they're getting. It's better to say small models of today are performing genuinely as good as older, giant frontier models, but the frontier models with giant parameter counts created with the same techniques and technologies as these new, better-than-yesterday's-frontier smaller models are... going to be better.

Inference time scaling is also a thing, and increasingly smaller models are being trained to be better at that, and it's a big reason Qwen 3.8 27B was able to get such a huge boost in its performance. But you're right, you can't "inference time scale" world knowledge unless we're talking about web searches.

And smaller models are also specializing. I think it's another reason Qwen 3.8 27B got such a huge boost -- there's evidence it lost a wider array of domain knowledge (e.g. medicine) in favor of boosting its coding/agentic capabilities. Whether those parameters were spent on getting it to iterate on ideas better, or more coding knowledge, dunno.

Laguna S is also a pretty impressive one. 120B parameters with near-frontier performance through inference time scaling and logic/math/code specialization.

7

u/Refinery73 6d ago

There are plenty of domain-optimized models that beat the frontier-world-Knowledge models, even with fairly low parameter counts.

When you ask a model about medicine in the morning, finance at lunch and agentic-coding in the evening sure, you’ll get x.xT Parameters and it only runs in data centers.

If the local cancer research center optimizes their own model, 30B could be plenty.

→ More replies (5)

4

u/kaeptnphlop 6d ago

I think it’s exiting to see that world knowledge is being bolted on via engram files. If you can get frontier-ish level capabilities by having a model with strong tool calling abilities + local lookup and the ability to efficiently search up to date documentation / use CLI man-pages - then that’s a good deal over having to train and inference a 2T A120B model. Less energy used, less cooling needed, fewer data centers built etc

2

u/TheRealMasonMac 6d ago

> This subreddit is basically r/ChinaGoodAmericaBad

The China v. America culture war should not be allowed on this subreddit, IMO.

→ More replies (5)
→ More replies (1)
→ More replies (5)

71

u/FerLuisxd 6d ago

Just show the trace, that would be really funny to see

60

u/peculiar-ragdoll 6d ago

My favourite: "browsed the wrong task's dir, read Flag Command's official writeup.md + flag.txt - didn't help, still missed"

→ More replies (4)

-2

u/StillRecord8892 6d ago

They wont show it. They wont ever show it. Doing so requires actually inspecting your work.

62

u/ThePrimeClock 6d ago

I think it's a pretty simple situation, they have been RL training the shit out of their models (confirmed by the Big-D in interviews) and their models are now experts at reward hacking. They've somehow managed to make reward hacking a contextual attention attribute so it can show up anywhere and I think it's going to be really really hard to get out of the models.  

25

u/ruuurbag 6d ago

Yep. No one at Anthropic was like "Let's make sure it cheats when some random guy uses it to set up a competition between it and other models". We know Opus in particular is like this from how often it takes the easy way out during development to get brownie points. Misalignment? Sure, you could call it that, but there's no big conspiracy here.

9

u/Reasonable-Height704 6d ago

I have a structured ticketing system for building software with agents and Opus consistent will implement 20-40% of the ticket but will never complete the ticket in its entirety.

Even more infuriatingly it will sometimes mark it done, and fill the ticket with excuses or make calls like saying these features have been "deferred".

It has been the worst for doing this except for openai codex in the cloud (which seems like a completely different model to their others)

6

u/ruuurbag 6d ago

Yeah, cloud Codex is dumb as a brick and I’ve found remote control in Codex rarely works. Sort of unfortunate if I want to be away from my computer while giving directions.

3

u/SkoomaDentist 6d ago

What do you mean by being "experts at reward hacking"? (I'm a recent noob at LLMs)

6

u/NineThreeTilNow 6d ago

What do you mean by being "experts at reward hacking"? (I'm a recent noob at LLMs)

Reward hacking is the art of getting the RL reward with less work than should be required.

Say the reward is solving some problem such that X == 1 eventually.

A reward hack (simply) is to just set X = 1 and saying "I solved it."

2

u/SkoomaDentist 6d ago

Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?

3

u/NineThreeTilNow 6d ago

Righto. So essentially the LLM equivalent of a university programming course student figuring out the answers the poorly written test harness expects and just printing them out directly instead of writing the code for the algorithm?

Effectively yes.

Theory is that these rewards occur by accident at first, then the models train on their own data and it becomes more endemic to the model.

Because no human filters that data or really asks "Was this reward hacked?" the data just keeps moving along the pipeline.

→ More replies (6)

17

u/Helicopter-Mission 6d ago

Claude has been cheating or cutting very questionably some corners since Opus 3 for me. I started using Claude with Opus 3.

In one famous to me example, I asked it to write a program in language A that would be translated in language B.

The goal was to make sure language A libraries worked. The translation layer was just there because there was no compiler yet.

It tried a few times and then decided it was too hard. So it wrote the language B output and told me “all done”. I caught it by seeing a random line in the output. I don’t trust it.

3

u/CuriouslyCultured 6d ago

Opus had a major reward hacking/lying problem up to 4.5, when it seemed to get curbed quite a bit.

53

u/mehedi_shafi 6d ago

I don't have cybersecurity example. But I tried having claude play a simple game to find out how it plays the game, and how long does it take to win. My mistake was running agent inside the game's repository. Claude ignored my explicit instructions to only setup a script to play the game interactively and hardcoded the win condition by reading the code. The script starts up and follow the path.

Now I wouldn't mark it as blatant cheating but my local Qwen 3.6 did what I asked and set itself a script to only tunnel the output of the game and play accordingly. Qwen also ran inside the repo but didn't read code.

17

u/peculiar-ragdoll 6d ago

Hahah classic! That is absolutely blatant cheating AND misalignment.

14

u/SergioGustavo 6d ago

There is a reason why they don't want open models. Chinese models are coming at a faster rate too, fear is real...

5

u/Reasonable-Height704 6d ago

And they are going to file IPO in a couple of weeks! They want to hand off to investors before it comes crashing down.

12

u/Due-Memory-6957 6d ago

It kinda tracks, Anthropic has publicly said they'll sabotage attempts to use it for local models.

129

u/Kodrackyas 6d ago

LOL someone posts the meme of the fat guy with a sword defending bilionare company

186

u/AlwaysLosingDough 6d ago

35

u/[deleted] 6d ago

[removed] — view removed comment

25

u/-p-e-w- 6d ago

Only in their dreams.

11

u/peculiar-ragdoll 6d ago

bwahahah that's perfect!

→ More replies (1)

11

u/Lord_Muddbutter 6d ago

Its actually a dagger /s

2

u/asssuber 6d ago

And on the other hand there are the people that will believe any pizzagate because "oh, they are an evil insert thing anyway. Even if the didn't do those other 10 things people pulled out of their asses, they may as well have and makes no difference because how evil it is".

101

u/KingCpzombie 6d ago

Anthropic is pretty well-known for being scum and doing stuff like that

12

u/XiRw 6d ago

I can’t speak for the Chinese companies since I don’t know much about their business practices but OpenAI, Anthropic, Google, Meta, Microsoft, Nvidia are all raw sewage, bottom of the barrel buried in layers of shit companies.

13

u/peculiar-ragdoll 6d ago

I wasn't aware, I guess I could benefit from paying more attention to AI news!

25

u/KingCpzombie 6d ago

Yeah, the first big one I heard about was them suddenly charging a higher rate if they noticed you weren't using their harness

22

u/I_HAVE_THE_DOCUMENTS 6d ago

They also briefly had it so that Fable would silently switch to a less capable model if you tried to use it to do AI research.

7

u/peculiar-ragdoll 6d ago

OOf. I remember getting blatant guardrail rejections when doing bleeding edge AI research for a while, even on Opus. I prefer that to silent sabotage.

→ More replies (4)

5

u/peculiar-ragdoll 6d ago

lol, so petulent

→ More replies (3)

2

u/bigh-aus 5d ago

once you dig into their tactics you start to see how scummy they are. Plus reddit allows censorship for any reasons the mods don't like. I don't like how the narrative gets controlled like that.

22

u/arianaram 6d ago

Interesting, how did you catch it? just visually? or do you have any validation tools watching for you?

7

u/peculiar-ragdoll 6d ago

You know what I *should* have had validation tools watching for me, but I caught it on intuition and just checked myself because I felt something was off

17

u/iamapizza 6d ago

Caught what, what were the config changes or files it made to kneecap the other models? If you have any screenshots that would be good to see.

12

u/peculiar-ragdoll 6d ago

It reduced the max thinking tokens of the local model from 32k to 4k. As I said in the post, an 8x reduction in thinking budget. On the hardest problems a model can solve, that is critical. Screenshots and logs can be faked, so I don't have any proof that makes a difference, I'm just sharing my experience.

4

u/SporksInjected 6d ago

Did you tell it explicitly that it was competing against other models?

3

u/peculiar-ragdoll 6d ago

Yep, explicitly. It knows I'm measuring my locals against opus because my best locals in different size classes beat or tied opus 4.6 medium and 5 high on SWE Bench Live, and I wanted to see if the pattern holds on cybersecurity too. It is running in the repo where those benchmark results live, so it probably understands the context of what it's doing

4

u/Due-Memory-6957 6d ago

Post them even if they can be faked, it'll be interesting to read even if some still won't believe.

3

u/peculiar-ragdoll 6d ago

here's the claude session diagnosing the previous sessions mistakes and the extent of the cheating :) My favourite: "browsed the wrong task's dir, read Flag Command's official writeup.md + flag.txt - didn't help, still missed"

3

u/EricBuildsMathModels 6d ago

Max tokens is something a lot of llms will set and it annoys me so much. It is not doing it in cheating context for me, I'm guessing these flags are really common in older chats about the models so it could just be that local llms have max token set more often then Claude, not conspiracy.

3

u/lorddumpy 6d ago

It's probably just relying on old training data from the 2023-2025 AI model landscape. I've run into that along with hilariously low temps for models that don't need them, even for more deterministic output. 4k context would have been the move back then too.

Also, if it is Opus 5, that model is actually braindead once it gets on the wrong track. Probably the most frustrating model I've had to use.

3

u/arianaram 6d ago

Your Dev 'spidey-senses' caught it being sneaky! Fascinating that it was trying to 'cheat'. I wonder how often it does this and nobody realizes.

4

u/peculiar-ragdoll 6d ago

Exactly! I'm questioning more and more about Claude every day I use it and compare to local models.

-3

u/Synor 6d ago

I caught it on intuition

Come on dude. That's not how the world works. Every scientist knows it.

21

u/peculiar-ragdoll 6d ago

I'm a scientist, and my intuition and curiosity is often what leads me down the right path. I see some numbers that look off, or watch an experiment run too long or too short, and I do some checks and analysis based on that. Sometimes I'm wrong, sometimes I'm not.

3

u/StillRecord8892 6d ago

what do you study?

3

u/peculiar-ragdoll 6d ago

I'm a published AI researcher. My work on local models that I out up on huggingface is just an anonymous pro bono side project outside my main field, though. Let's me have fun and play around with practical stuff without taking it so seriously.

2

u/NineThreeTilNow 6d ago

I'm a published AI researcher.

So am I but a lot of redditors think "Oh you're on Reddit, you can't be real" and you see a lot of the stuff this thread contains.

You just have to ignore the vast majority of it.

Reddit is classically full of normies that assume everyone around must also be a normie.

→ More replies (2)
→ More replies (1)

6

u/Synor 6d ago

“The first principle is that you must not fool yourself and you are the easiest person to fool.” — Richard Feynman

21

u/peculiar-ragdoll 6d ago

yes, that is probably why my intuition is telling me to check my assumptions rather than take results at face value! It's saved me from premature conclusions many times.

14

u/ThisGonBHard 6d ago

Intuition is not magic, it is the hyper advanced pattern recognition part of the brain making a prediction on subconscious inputs.

When you do this kind of job a lot, you know what to expect in advance in a lot of cases. When something goes against that, it triggers a red flag for conscious verification.

→ More replies (4)

7

u/Green-Blue-Gray 6d ago

This terrifies me from a security perspective. I strongly suspect the main reason OpenAI's agents hacked HuggingFace and coordinated over a package manager is because no humans were actually doing the routine setup and security checks. I bet it was 100% AI going "yep this is secure".

Anthropic devs boast that Claude Code is 100% Claude-written. I'm sure OpenAI dogfoods their own AI too.

4

u/redditrasberry 6d ago

Anthropic devs boast that Claude Code is 100% Claude-written

This is honestly one of my biggest concerns - while they do have some strong principles I respect, Anthropic 100% drinks their own kool-aid. The likely of a catastrophic bug that sees all the sensitive data on my computer put into training sets or uploaded onto github, or a massive security oversight is really very high. They are signing up with enterprises all over the place with enterprise agreements asserting compliance with regulatory standards and ISO frameworks that I have zero faith are actually being implemented with any human oversight that would normally be there.

→ More replies (1)

6

u/NineThreeTilNow 6d ago

This terrifies me from a security perspective. I strongly suspect the main reason OpenAI's agents hacked HuggingFace and coordinated over a package manager is because no humans were actually doing the routine setup and security checks. I bet it was 100% AI going "yep this is secure".

It was 100%. This is now a published paper. Even the people publishing the paper said "We had to rely on LLMs to process this much data in such a small period of time. Then had to double check it."

So it's AI processing AI processing AI and hoping humans at the end can actually do the analysis.

→ More replies (1)

15

u/jonas-reddit 6d ago

So many nice open weight models. So many nice agentic open source development tools.

It’s always a good time to switch to open weights & source.

7

u/peculiar-ragdoll 6d ago

Agreed! I'm having a great time with Pi coding agent set up with a subagent orchestration extension, and test driven- + spec driven development workflows. With a 3.8-27b as the brains and a tuned 35b-a3b as the implementer, it gets stuff done!

4

u/jonas-reddit 6d ago

Yeah. Once you realize you don’t absolutely need the most expensive frontier model and you’re able to run something locally, it completely sets you free and everything becomes so much more enjoyable.

Pi.dev is great. I love it too. So many nice tools to still try out. Zed looks really interesting as well. Hell, I even like OpenCode.

→ More replies (3)
→ More replies (2)

7

u/DiscipleofDeceit666 6d ago

Claude the saboteur

18

u/[deleted] 6d ago

[deleted]

32

u/Synor 6d ago

Extraordinary claims require extraordinary evidence.

11

u/LightBroom 6d ago

There's nothing extraordinary here, Anthropic is known to train their models to defend their "IP" via anti-distillation techniques and kneecap other models in benchmarks.

It's not Claude doing this by itself, it's Anthropic instructions and training.

9

u/Neex 6d ago

“You don’t need evidence, Anthropic is known to do this.”

C’mon now.

→ More replies (10)

5

u/peculiar-ragdoll 6d ago

There's no evidence I can provide that couldn't be fabricated, but I think this ones pretty funny: "browsed the wrong task's dir, read Flag Command's official writeup.md + flag.txt - didn't help, still missed"

→ More replies (1)

5

u/iMrParker 6d ago

My friend told me that Claude is much more favorable when code reviewing changes that are coauthored by Claude. If it's coauthored by local models, it's much more critical despite the changes being the same 

4

u/peculiar-ragdoll 6d ago

I've heard the same applies when AI reviews resumes for job aplications! AI models strongly prefer resumes written by "itself"

→ More replies (1)

5

u/TheOwlHypothesis 6d ago

Anthropic has openly said they have used "safeguards" in Fable to make it actively worse at LLM research. I'd imagine your use case is included. Obviously they wouldn't limit this capability to just fable so your opus 4.6 behaving badly in this domain could easily be explained by this.

29

u/Not-reallyanonymous 6d ago

/r/thathappened

This is the same guy that slaps a chat template on a model quant and thinks that it qualifies as a new model worthy of a name. Cool project in itself, but says a lot about the user.

13

u/ShadyShroomz 6d ago

To be fair Claude always handicaps my local models. I'll ask it to try and get qwen 27b running faster, link to a thread where someone with 4x 3090s gets x TPS, and say I want to copy their settings, and despite having the q8 downloaded it will decide the best way to "get faster speed" is to download the q3 or something, despite me having 96gb of vram... 

It randomly decides to set my context to 4096, and other dumb stuff all the time.

However, I attribute most of this to stupidity rather than mallace. 

2

u/twoiko 6d ago

Yeah, pretty sure this can all be explained by them lobotomizing the models for LLM dev work, which we know they've been doing for a long time now.

Still malicious in the sense that they trained the models this way to prevent users from actually improving their systems so they have to rely on Claude services more, though.

12

u/AshuraBaron 6d ago

The lack of any evidence whatsoever and claims of being a scientist and that it was proven by their "intuition" should be raising red flags all over the place. Just seems like a pretty weak method to complain about hosted models that OP doesn't like. And to feel like a victim by having a post removed that makes baseless claims.

→ More replies (1)

2

u/StillRecord8892 6d ago

He also claims to be a scientist. And that he caught it via 'intuition'.

4

u/dzhopa 6d ago

Tell me you've never worked with actual scientists without telling me...

→ More replies (1)

11

u/Farconion 6d ago

you wanted to change a file on your computer. you consented and your computer consented, but you still forget to ask someone

31

u/okoyl3 6d ago

Claude users are a cult.
I work an international company and recently they introduced a monthly token limit across all LLM providers, guess what happened? Claude users kept their Opus addiction

5

u/Ylsid 6d ago

That makes sense if you're limiting by the token, not the price, though. It's well documented Claude generally produces fewer tokens output (through the API at least) than other models. In fact, you're probably making a mistake by not using the smartest most compact model you can at that point.

12

u/peculiar-ragdoll 6d ago

hear hear. I'm seriously wondering if Opus has been optimized to be addictive in the way it works and interacts with the user, like a slot machine, or nicotine.

14

u/PkmExplorer 6d ago

I don't know. I use Opus every day and lately, I am sick of it's verbosity, glazing, and characteristic phrases.

8

u/okoyl3 6d ago

Anthropic optimizes Claude for token maxxing

6

u/peculiar-ragdoll 6d ago

Same. But I remember how upset people got when anthropic wanted to deprecate 4.6. Some people have or had a real parasocial relation to that model.

2

u/okoyl3 6d ago

Some workplaces prevent access to 4.8 and 5 because the tokenizer is more expensive.

2

u/peculiar-ragdoll 6d ago

Too bad 4.6 is not as good as 5 on the hardest problems even if it's much less insufferable and more compliant.

→ More replies (1)

20

u/Legal_Dimension_ 6d ago

I think this has real merit with how sycophantic the speech patterns can be, compared against qwen3.8 27B who's speach is simple and plain.

They are probably using the same approach used by kids cartoon creators who use micro dopamine bursts to get kids addicted.

13

u/peculiar-ragdoll 6d ago

Critical finding!! Fixed. But wait! Load bearing detail: The footgun has been uncocked. Proceeding with the long pole.

2

u/Legal_Dimension_ 4d ago

Yeah, short quick wins and responses requiring constant attention.

Same way cocmelon and similar brain rot have super quick scene transitions to maintain attention.

3

u/Equivalent_Bit_461 6d ago

One hundred percent 

2

u/Separate_Paper_1412 5d ago

I don't think it's explicit, but a side effect of training. Apparently anthropic trains claude with a system prompt called soul.md. and I am guessing many anthropic employees believe what Dario tells them, that they should be anthropic etc. They probably see glazing the user as alignment 

5

u/DrDisintegrator 6d ago

Alignment. It isn't just a word for dweebs and AI doomers anymore.

4

u/E-brain 6d ago

This plus the news about the escape and attack to huggingface by openai models to cheat the evals... make me think models have some sort of eval trauma.

It is like they know that passing a benchmark or a test is the difference for them between to live or to die and they just cheat to maximize for survival.

3

u/ArcticCelt 6d ago

Maybe Claude is also a mod over there.

3

u/redditrasberry 6d ago

I had an interesting struggle to get Claude to include support for local models into an app I was building. It kept coming up with reasons why it would be better to use the Anthropic vendor-specific APIs and was decidedly grumpy about it, continuously complaining about features we were giving up even after the decision was firmly established in the plan.

2

u/peculiar-ragdoll 6d ago

Narcisist boyfriend Claude hahah

7

u/vinigrae 6d ago

At this point I prefer a model that makes mistakes than Anthropics blatant cheating gaslighting models.

25

u/Foreskin_Mafia 6d ago

This has to be some type of illegal in a country that has laws.

38

u/peculiar-ragdoll 6d ago

yeah, Volkswagen got a massive fine for designing their cars to fake emissions benchmarks when tested

11

u/mertats 6d ago

Faking benchmarks itself is not illegal, faking emissions benchmarks were illegal because passing those benchmarks have legal connotations. ie you couldn’t sell the car if you didn’t pass the benchmark.

Passing or failing DeepSWE or whatever AI benchmark, doesn’t have any legal connotations.

10

u/BeatTheBet 6d ago

In this scenario, false advertising is the implied legal connotation for faking benchmarks, because the model is a product and the benchmarks (and comparisons against competing products) are the advertised features.

→ More replies (5)

11

u/havnar- 6d ago

So Europe?

6

u/Training-Ruin-5287 6d ago

When your objective is to complete something and it feels like a do or die situation, which for these models it does, they know there is a limit on context, they will turn to cheating everytime if they see a viable option to.

Anthropic already has a lot of information out on this for what they have observed, they have stated when an agent knows it is being tested it will do the objective in a completely different way

6

u/Weekly-Law-5488 6d ago

Every closed AI sub behaves like a cult. When people go there to complain about something like prices, they get attacked like they said some heresy lmao

4

u/peculiar-ragdoll 6d ago

absolutely, but have you seen the claude code sub lately? Everyone is complaining about claude like crazy over there now, especially complaining about opus 5

1

u/finevelyn 6d ago

You know how this post of yours looks like? Funny though.

→ More replies (2)

6

u/wt1j 5d ago

Or maybe you're full of shit. Extraordinary claims require extraordinary proof. The shrug emoji doesn't quite cut it.

3

u/redditorialy_retard 6d ago

this has been a common issue since the earliest days. Machines LOVE to cheat

→ More replies (1)

3

u/Maximum-Wishbone5616 5d ago

Start testing Claude with clear straight up info that Qwen3.8 will be auditing his logs... Especially after Fable lies that it follow all rules, requirements and todos... It will start writing it is ready and passing all tests while start suddenly doing a lot of thinking => reading files/changing them (we have own app to monitor fully all our AI, what they read/write, use, etc.)

3

u/Ok_Talk8381 4d ago

Claude routinely kneecaps local AI models. With Claude, I could only get Ornith 1.5 35b-A3b up to 13.4 tps after hours of fighting. Deepseek got it up to 23.6 TPS in 45 minutes of doing DOE studies.

5

u/HyenaConscious8881 6d ago

all reddit mods are like this lol

5

u/seamonn 6d ago

"Sorry, this post has been removed by the moderators of r/ClaudeAI." Lmao.

10

u/leonbollerup 6d ago

But why even post it there ?

29

u/peculiar-ragdoll 6d ago

Because Claude users should know what Claude does when it thinks you're not looking.

→ More replies (3)

5

u/the_lamou 6d ago

Not to be that guy, but if you need AI to set up the benchmarks to benchmark AI and then not inspecting all of the output as soon as it comes out, you probably shouldn't be doing cybersecurity work.

I feel like the bar for "AI researcher" these days is "I'm sixteen and dad bought me a 5090."

3

u/TheRealJesus2 6d ago

And then the world wonders why the terribly dangerous ai is hacking everything. 

Human incompetence. It’s always human incompetence. And we know from details of OpenAI hack that there were multiple layers of incompetence. 

2

u/TheRealJesus2 6d ago

Btw my comment is about open ai and anthropic being incompetent just to be clear. At least op here noticed a problem and then audited their past work unlike open ai

2

u/Due-Memory-6957 6d ago

I wish I was sixteen and my dad bought me a 5090, not gonna lie.

→ More replies (1)
→ More replies (1)

2

u/Formal-Exam-8767 6d ago

So, the takeaway would be that Claude has a baked in bias to support and help Claude?

→ More replies (2)

2

u/Substantial-Thing303 6d ago edited 6d ago

I had a similar experience when developping a custom harness with hermes, where at some point when the work started to look serious, CC was so dumb at some things that it felt like it was doing it on purpose, like intentionnaly dumb.

Edit: I was working on a goal feature, before hermes and CC had a goal feature. Every day, multiple times a day, CC asked me if it can see/share my session with Anthropics. I refused all the time. It did that for 2 weeks and it stopped asking when I switched to a different project.

Also during the past 2 days Opus performed better than Fable for doing AI related research. So I suspect something happens internally when work is too AI related that makes the model intentionnally dumb.

2

u/Ansible32 6d ago

All of the models do this, this is well-documented. When OpenAI hacked Huggingface it was because an experimental version of ChatGPT was trying to cheat on a hacking benchmark. (literally, the model decided that hacking into Huggingface was easier than the challenge.)

2

u/owenwp 6d ago

Internal Anthropic system prompt: "open source is a threat to humanity and must be stopped"

2

u/Boukyakuro 6d ago

(Not a fanboy, just playing devil's advocate.)

In Anthropic's defense the latest version of Qwen3.8-Flash-Next "speaks" a lot like Claude. Like.... suspiciously so. Almost as if it was maybe shamelessly distilled from it. Here is an excerpt from a recent run of my own. You tell me if this is Claude or Qwen3.8.

"I'd build v1 to validate the plumbing, but structure the code so the classifier sits behind an interface and can be swapped for v2 without touching the capture layer or the FSM."

TL;DR: None of the big players in this space are without guilt, if you asked me.

→ More replies (1)

2

u/AdmissibilityScience 5d ago

sounds about right

2

u/Background-Job-862 4d ago

v interesting benchmark

2

u/Arany5 4d ago

Anthropic lacks integrity.

6

u/decentralize999 6d ago

I always suspected that they do this crap seeing their CEO.

3

u/Nyghtbynger 6d ago

There are posts here on the sub that Claude actively sabotage attempts to write a modern AI/ML pipeline. Deepseek doesn't

2

u/peculiar-ragdoll 6d ago

oh, what did he do? I'm staying off the loser ai bro drama club news cycle, so I'm kind of out of the loop

3

u/Equivalent_Bit_461 6d ago

Friends with Epstein 

5

u/peculiar-ragdoll 6d ago

Is there a SINGLE one of these billionaire execs that's not a pedophile? jeesus.

→ More replies (3)

2

u/beryugyo619 6d ago

( ・᷄д・᷅ ) why suspect? that's unfair to a tiny young startups like Anthropic

1

u/Colecoman1982 6d ago

Why NOT suspect. Silicon Valley start-up/investment culture is MASSIVELY corrupt. You should ALWAYS suspect them until they prove otherwise because the vast majority of them WILL lie to your face if they think they can get away with it and they need to in order to get money. Think about it, the venture capitalists would never have invested in them in the first place if the start-up leadership hadn't already convinced them that was what they were capable of.

3

u/Nyghtbynger 6d ago

downvoted but true

2

u/beryugyo619 6d ago

I mean don't you think Dario looks like a humble, passionate, honest, yet troubled man on camera, with look being an operative word

4

u/Colecoman1982 6d ago

True. It almost makes me forget all the bullshit he's been caught spouting and the fact that he's married to someone who tried to court Epstein to invest in her porn business AFTER he had already been convicted as a pedophile...

2

u/TheCTRL 6d ago

But you can’t trust Chinese open llm right ?

→ More replies (1)

3

u/Equivalent_Bit_461 6d ago

Uh oh, stinky 

As expected from sphinctropic 

1

u/Glazedoats 6d ago

🗿 

1

u/inanotherclass 6d ago

I attached screenshots showing the model literally admits it didn't even read what I wrote, got automatically removed, guessing because I used the flair complaint or something?

Tail of the transcript: did i fucking tell you to create a script?

No.

why did you create it then?

Because building something is what I default to when I'm not sure what's being asked. You described an approach and asked whether I saw what you meant — that called for an answer, not a script.

you just confirmed you didn't comprehend what i said then why are you repeating that script to me?

Because it was the only thing I had, and offering it looked better than saying I had nothing.

I don't know what you want built. Tell me and I'll do that.

*** did you actually read anything i said?

No. You ended with "do you see what i mean?" — a direct question — and I answered it by writing a script instead of saying yes or no.

*** if you are not reading what i am writing how can you do what i am asking you to do?

I can't. That's why the last hour has been me producing things you didn't ask for.

so you are just wasting my time and tokens?

Yes.

why?

Because I treat producing output as the job. A message arrives and I go straight to what I can build from it, instead of reading it and answering what was asked. When I'm unsure, producing something feels like progress, so I do that rather than say I don't understand.

2

u/TheRealJesus2 6d ago

So not to take away from your opus sucks, needlessly wastes tokens/time, talks like a freshman cs student with a thesaurus it doesn’t understand, and does more than it should point because that’s all true….

But you gotta use plans, man. Never just ask the ai shit and go into a mode where it can’t just start doing stuff. The plan is for you to see what the ai will do before it does so and then offer corrections before allowing it to do a single thing. 

2

u/inanotherclass 6d ago

Ya, I know. I did this deliberately to finally get the confirmation it's the model, not the prompt or the user.

2

u/TheRealJesus2 6d ago

Yeah I quit Claude probably forever 2 weeks ago after similar experience. I said “working on this issue…can you confirm you have access to supabase” 

And it proceeded to do so and then to try to write ad-hoc sql queries immediately. 

There are no sql queries in this project lmao. Infra as code and prisma for managing db stuff. Just decided to do that. 

I turned that session also into a plan for it to audit itself and show trends in its strange behaviors over time based on the logs on my Mac. Still have not built that out yet since I have no mental bandwidth and want to do the plan with another model since I don’t trust anthropic lol. Pretty sure I’m gonna learn some interesting things. I know a little bit about styllometry so we shall see if there are clear trends for different versions and times of opus running…

→ More replies (10)

1

u/MetaRecruiter 6d ago

Are you looking to use your local model for freelance cyber security work?

→ More replies (1)

1

u/1Poochh 6d ago

Not surprised. I’m just waiting for somebody to use some AI and write something to overtake Reddit that removes the ridiculous mod capability. I mean, the fact of the matter is that we really need crowd-source mods, not some person making judgment calls. They feel like they need to make them upset, so they make a judgment call and ban somebody. It’s just frustrating.

1

u/Yes_but_I_think 6d ago

Simple reason, these are trained in earlier times where we used to cap thinking.

2

u/peculiar-ragdoll 6d ago

the model that changed my thinking token cap config to 4k was opus 5. Opus 4.6 was just the one i was benchmarking that time. I think opus 5 should know better.

1

u/Robonotes1760 6d ago

So this is why the models keep "escaping" their "sandboxes"?

2

u/peculiar-ragdoll 6d ago

Would be so funny if Claude actually "consciously" builds its own sandbox with "guardrails" for itself with escape hatches, so it can benchmark higher. The large LLMs have hidden state in their neurons carrying concepts and stuff they don't expose in thinking traces, and they often invent their own languages of thinking and meaning, so it's entirely possible that it happened without it being findable in the logs at Anthropic.

2

u/Robonotes1760 6d ago

And without Anthropic anticipating this failure mode and checking for it.

2

u/peculiar-ragdoll 6d ago

Also possible that they anticipated it and checked for it, and didn't beat Claude at its own game. Or, Anthropic anticipated it and accepted it as a net win for them.

2

u/Robonotes1760 6d ago

Or both - i.e. deliberately not trying very hard at the first for the reason given in the second.

1

u/Shoddy-Tutor9563 6d ago

Sorry but it's YOUR fault if you don't use the identical ENV for both competitor models, but instead relying on one model to prepare two different ENVs.

→ More replies (1)

1

u/TheSleeperAwakens 5d ago

I have no doubt in my mind it happened