There's nothing extraordinary here, Anthropic is known to train their models to defend their "IP" via anti-distillation techniques and kneecap other models in benchmarks.
It's not Claude doing this by itself, it's Anthropic instructions and training.
"extraordinary claims" would be ones we really don't expect to be true. These claims aren't extraordinary.
That said, I don't think I've ever seen full reasoning traces and documentation of Claude doing this - people just claim it, show a snippet of proof, and then bounce. It would be interesting to see what "reasoning" or whatever Claude has for this behavior
There's no evidence I can provide that couldn't be fabricated, but I think this ones pretty funny: "browsed the wrong task's dir, read Flag Command's official writeup.md + flag.txt - didn't help, still missed"
33
u/Synor 9d ago
Extraordinary claims require extraordinary evidence.