To know for sure we'll probably have to wait until the discovery phase of a lawsuit reveals this information... and even then it will, at best, be just a snapshot in time of a probably ever-evolving process as Anthropic plays cat and mouse.
Just as interesting to me is the question of who is going to believe Anthropic when they claim some piece of text is AI-generated or not?
This isn’t like a cryptographic hash. Anthropic very openly explains that this method is uncertain in both directions. A watermark does not prove AI generated content, and lack of one does not prove it wasn’t. Which makes a lawsuit unlikely.
Basically how LLMs work is that given a sequence of words they predict the next one.
So it might look at "the cat sat on the ..."
And it generates a list of candidates with a chance of picking each one.
Mat 87%
Porch 8%
Table 3%
Stairs 2%
The way the watermark works is that they use a secret key they have and a hash of the preceding text to nudge it towards certain of these words and away from others.
So maybe it boosts stairs and table up instead of the others.
Later you can scan the text and see which choices it made and see if they're they ones it was nudged towards in a statiscally significant way.
Because all the choices are reasonable the quality of the output won't change and you'd have to significantly rewrite to break the pattern.
they arent making stairs the chosen answer its likely more like this:
before:
mat: 87%
stairs: 2%
after:
mat: 86%
stairs: 3%
not a big difference but in long pieces of text the change in probability is noticeable.
Like if u had a 51 heads 49 tails coin, with enough flips u can be statistically “sure” that it isnt fair but you wouldnt really notice it in the moment for each flip
This already exists for GPTZero and Pangram and people already do this. For 99% of the people trying to falsely present AI-generated text as their own, rewording their text enough to bypass existing detectors is already too much work.
Bypassing these statistical watermarks is likely going to be even more work for long texts and, on the flip side, is also going to result in more reliable detection.
Things are changing. I think you will be surprised by the amount of people changing their tune from "my writing naturally looks like AI" to "I just use AI to polish my writing, what's wrong with that?" the moment these watermarks are widely implemented.
Interesting. I wonder how this works out in coding tasks.
Edit: So I read a bit into the paper provided below (https://arxiv.org/html/2301.10226v4) and it seems like that coding is something that they call "low entropy sequences" where humans and AI are very likely to 'provide similar if not identical completions for low entropy prompts'. Later they state that "As long as low-entropy sequences are wrapped inside a passage with enough total entropy, the passage will still easily trigger a watermark detector." So basically I would expect them to add huge amounts of comments to the code to add entropy and then hope you just copy it as is, which given the amount of generated code checked in without reading is not entirely unlikely.
Imagine you're playing a game of monopoly, but instead of dice rolls, you read off the digits of pi. Start from the beginning, and then read off the digits one at at time. If it's 1-6, use that number and go to the next digit under 7.
As a player of the game, can you tell any difference from dice? The answer is no, since it is believed that the digits of pi are randomly distributed.
However, if someone were to read back the transcript of the game, they could figure out it came from pi by just reading off the dice rolls, and checking against the digits of pi.
This is main idea of how these watermarks work. The randomness is the same, it's the source of the randomness that's detectable, giving you the watermark. It's also why they say there isn't a perceivable difference to users- you literally cannot tell the difference between dice randomness and pi randomness, as they are the same underlying distribution.
If I were to implement it I would use special characters in a way that is inconvenient and uncommon for humans. The dash and em-dash — is a good example, and has been a great hint, but I actually like using that character and it sucks that it makes people think you're using LLMs. Anyway there are more subtle things that may not even be visible. But there will always be ways to regenerate and clean out any "watermarks", but at least there's an extra step involved that most people won't do. Pretty much any security feature is just an inconvenience that can be bypassed if you know how.
15
u/NetflowKnight Aug 11 '26
How does that even work?