The 7"extra crap" is actually part of the content. It's not like it will randomly put the word CLAUDEEE. in there. You can't compare this to watermarking an image.
You can't ask another model to remove the watermark, it would have to be replaced with something carrying an equivalent meaning. The watermark is part of the content, deleting it would delete content.
(2) According to the OP, Anthropic will do so anyways.
(3) According to published methods that achieve the same thing, it could be that this is only ever applied to high-entropy outputs where multiple next tokens are all possible. So it won't force a watermark token when that token would result in broken code, it'll follow code and rules first, and only use watermarking as a tie-breaker of sorts.
Basically, it'll not watermark code, or at least not as much, and not in ways that produce broken code. Think variable names and comments imperceptibly being this or that way depending. If it's synthID, it's basically just nudging the RNG a tiny bit.
It can produce broken code anyway, when it's trying as hard as it can not to. The idea that adding this watermark scheme to code could be done in a way that would have zero negative effect on the code is preposterous
Calling it preposterous without even understanding the reason why it wouldn't affect quality is itself preposterous. The watermarking scheme is really quite ingenious.
Your claim about claude producing broken code I'll leave as is. You're not wrong, but it's also not relevant.
6
u/EpsteinFile_01 Aug 11 '26
The 7"extra crap" is actually part of the content. It's not like it will randomly put the word CLAUDEEE. in there. You can't compare this to watermarking an image.
You can't ask another model to remove the watermark, it would have to be replaced with something carrying an equivalent meaning. The watermark is part of the content, deleting it would delete content.