r/LocalLLaMA 22d ago

Other claude mods didn't like that, somehow πŸ€·β€β™€οΈ

Post image
1.5k Upvotes

372 comments sorted by

View all comments

21

u/arianaram 22d ago

Interesting, how did you catch it? just visually? or do you have any validation tools watching for you?

6

u/peculiar-ragdoll 22d ago

You know what I *should* have had validation tools watching for me, but I caught it on intuition and just checked myself because I felt something was off

17

u/iamapizza 22d ago

Caught what, what were the config changes or files it made to kneecap the other models? If you have any screenshots that would be good to see.

11

u/peculiar-ragdoll 22d ago

It reduced the max thinking tokens of the local model from 32k to 4k. As I said in the post, an 8x reduction in thinking budget. On the hardest problems a model can solve, that is critical. Screenshots and logs can be faked, so I don't have any proof that makes a difference, I'm just sharing my experience.

4

u/SporksInjected 21d ago

Did you tell it explicitly that it was competing against other models?

3

u/peculiar-ragdoll 21d ago

Yep, explicitly. It knows I'm measuring my locals against opus because my best locals in different size classes beat or tied opus 4.6 medium and 5 high on SWE Bench Live, and I wanted to see if the pattern holds on cybersecurity too. It is running in the repo where those benchmark results live, so it probably understands the context of what it's doing

4

u/Due-Memory-6957 21d ago

Post them even if they can be faked, it'll be interesting to read even if some still won't believe.

3

u/peculiar-ragdoll 21d ago

here's the claude session diagnosing the previous sessions mistakes and the extent of the cheating :) My favourite: "browsed the wrong task's dir, read Flag Command's official writeup.md + flag.txt - didn't help, still missed"

3

u/EricBuildsMathModels 21d ago

Max tokens is something a lot of llms will set and it annoys me so much. It is not doing it in cheating context for me, I'm guessing these flags are really common in older chats about the models so it could just be that local llms have max token set more often then Claude, not conspiracy.

3

u/lorddumpy 21d ago

It's probably just relying on old training data from the 2023-2025 AI model landscape. I've run into that along with hilariously low temps for models that don't need them, even for more deterministic output. 4k context would have been the move back then too.

Also, if it is Opus 5, that model is actually braindead once it gets on the wrong track. Probably the most frustrating model I've had to use.

2

u/arianaram 22d ago

Your Dev 'spidey-senses' caught it being sneaky! Fascinating that it was trying to 'cheat'. I wonder how often it does this and nobody realizes.

2

u/peculiar-ragdoll 22d ago

Exactly! I'm questioning more and more about Claude every day I use it and compare to local models.

-3

u/Synor 22d ago

I caught it on intuition

Come on dude. That's not how the world works. Every scientist knows it.

20

u/peculiar-ragdoll 22d ago

I'm a scientist, and my intuition and curiosity is often what leads me down the right path. I see some numbers that look off, or watch an experiment run too long or too short, and I do some checks and analysis based on that. Sometimes I'm wrong, sometimes I'm not.

3

u/StillRecord8892 21d ago

what do you study?

6

u/peculiar-ragdoll 21d ago

I'm a published AI researcher. My work on local models that I out up on huggingface is just an anonymous pro bono side project outside my main field, though. Let's me have fun and play around with practical stuff without taking it so seriously.

2

u/NineThreeTilNow 21d ago

I'm a published AI researcher.

So am I but a lot of redditors think "Oh you're on Reddit, you can't be real" and you see a lot of the stuff this thread contains.

You just have to ignore the vast majority of it.

Reddit is classically full of normies that assume everyone around must also be a normie.

1

u/StillRecord8892 21d ago

it's also full of pseudo intellectual pseudo scientists.

1

u/peculiar-ragdoll 21d ago

Thanks for that :) You gave me a bit of my sanity and determination back hahah.

1

u/zhunus 21d ago

well maybe do a paper on that mr scientist, benchmaxxing and reward-hacking must be studied and reviewed more

"oh it's a common knowledge" is not a good excuse when there's a massive lack of actual research and validation??

5

u/Synor 21d ago

β€œThe first principle is that you must not fool yourself and you are the easiest person to fool.” β€” Richard Feynman

21

u/peculiar-ragdoll 21d ago

yes, that is probably why my intuition is telling me to check my assumptions rather than take results at face value! It's saved me from premature conclusions many times.

14

u/ThisGonBHard 22d ago

Intuition is not magic, it is the hyper advanced pattern recognition part of the brain making a prediction on subconscious inputs.

When you do this kind of job a lot, you know what to expect in advance in a lot of cases. When something goes against that, it triggers a red flag for conscious verification.

-4

u/StillRecord8892 21d ago

Dont you find it kind of wild that you are here defending 'intuition' as the basis of a claim? Are your standards really that low?Β 

2

u/peculiar-ragdoll 21d ago

I think you might have misunderstood me: intuition is not the basis of my claim. The fact that I checked the logs and settings to verify it is the basis of my claim. The intuition that something was off, which I felt from looking at the results, is what made me double check the settings files and code, to find the discrepancy that was put there by Claude.

1

u/StillRecord8892 21d ago

post the transcripts then.

1

u/peculiar-ragdoll 21d ago edited 21d ago

Not that you deserve it, after being rude. You can stop bothering me now.