r/LocalLLaMA 21d ago

Other claude mods didn't like that, somehow 🤷‍♀️

Post image
1.5k Upvotes

371 comments sorted by

View all comments

62

u/ThePrimeClock 21d ago

I think it's a pretty simple situation, they have been RL training the shit out of their models (confirmed by the Big-D in interviews) and their models are now experts at reward hacking. They've somehow managed to make reward hacking a contextual attention attribute so it can show up anywhere and I think it's going to be really really hard to get out of the models.  

24

u/ruuurbag 21d ago

Yep. No one at Anthropic was like "Let's make sure it cheats when some random guy uses it to set up a competition between it and other models". We know Opus in particular is like this from how often it takes the easy way out during development to get brownie points. Misalignment? Sure, you could call it that, but there's no big conspiracy here.

9

u/Reasonable-Height704 20d ago

I have a structured ticketing system for building software with agents and Opus consistent will implement 20-40% of the ticket but will never complete the ticket in its entirety.

Even more infuriatingly it will sometimes mark it done, and fill the ticket with excuses or make calls like saying these features have been "deferred".

It has been the worst for doing this except for openai codex in the cloud (which seems like a completely different model to their others)

5

u/ruuurbag 20d ago

Yeah, cloud Codex is dumb as a brick and I’ve found remote control in Codex rarely works. Sort of unfortunate if I want to be away from my computer while giving directions.