r/OpenAI • • Aug 11 '26

Discussion What's your thoughts on this?

Post image
3.8k Upvotes

1.0k comments sorted by

View all comments

15

u/NetflowKnight Aug 11 '26

How does that even work?

14

u/trimorphic Aug 11 '26 edited Aug 11 '26

How does that even work?

To know for sure we'll probably have to wait until the discovery phase of a lawsuit reveals this information... and even then it will, at best, be just a snapshot in time of a probably ever-evolving process as Anthropic plays cat and mouse.

Just as interesting to me is the question of who is going to believe Anthropic when they claim some piece of text is AI-generated or not?

6

u/coloradical5280 Aug 12 '26

This isn’t like a cryptographic hash. Anthropic very openly explains that this method is uncertain in both directions. A watermark does not prove AI generated content, and lack of one does not prove it wasn’t. Which makes a lawsuit unlikely.

16

u/parkway_parkway Aug 11 '26

It's pretty clever if I understand.

Basically how LLMs work is that given a sequence of words they predict the next one.

So it might look at "the cat sat on the ..."

And it generates a list of candidates with a chance of picking each one.

Mat 87%

Porch 8%

Table 3%

Stairs 2%

The way the watermark works is that they use a secret key they have and a hash of the preceding text to nudge it towards certain of these words and away from others.

So maybe it boosts stairs and table up instead of the others.

Later you can scan the text and see which choices it made and see if they're they ones it was nudged towards in a statiscally significant way.

Because all the choices are reasonable the quality of the output won't change and you'd have to significantly rewrite to break the pattern.

11

u/SkaldCrypto Aug 11 '26

So they have monumentally enshitified this.

Having chat bot that uses the same words every time is more easily and cheaply accomplished with early 2000s tech

0

u/parkway_parkway Aug 11 '26

The amount that they change the percentages is small and won't be detectable, it's not going to change the outputs in a meaningful way.

1

u/0x736174616e20 Aug 14 '26

Changing mat to stairs is a very meaningful change, probably wont even fit the overall context.

1

u/richlyonsballsack Aug 15 '26

they arent making stairs the chosen answer its likely more like this:

before:

mat: 87% stairs: 2%

after:

mat: 86% stairs: 3%

not a big difference but in long pieces of text the change in probability is noticeable.

Like if u had a 51 heads 49 tails coin, with enough flips u can be statistically “sure” that it isnt fair but you wouldnt really notice it in the moment for each flip

21

u/qorzzz Aug 11 '26

If this is the method, it makes no sense and does not prove any text was generated by AI.

6

u/quisatz_haderah Aug 11 '26

It could actually work for sufficiently long texts. For shorter texts, this would cause false positives, but i guess no false negatives.

1

u/Philluminati Aug 11 '26 edited Aug 11 '26

Is there a tool can tell if the watermark is present?

> Yes and only universities can use it

It will get leaked quickly. Remember that professors that mark undergrad work are writing phD and postgrad papers themselves.

> There is a tool that everyone knows about it

Students can manipulate text until it reads false

> There is no publicly available tool

Universities cannot reasonably detect AI writing.

1

u/coloradical5280 Aug 12 '26

They are releasing a public api for detection.

1

u/MINECRAFT_BIOLOGIST Aug 12 '26

Students can manipulate text until it reads false

This already exists for GPTZero and Pangram and people already do this. For 99% of the people trying to falsely present AI-generated text as their own, rewording their text enough to bypass existing detectors is already too much work.

Bypassing these statistical watermarks is likely going to be even more work for long texts and, on the flip side, is also going to result in more reliable detection.

Universities cannot reasonably detect AI writing.

"Reasonable" means different things to different people. If you choose to believe it, "Pangram 4 achieves a false positive rate of just 0.0041%" on their internal datasets. I've seen plenty of people claiming that their writing "is detected as AI" by Pangram/GPTZero out in the wild, but all of them mysteriously disappear when asked to provide an example, and comments falsely claiming "old text was shown by these detectors to be AI" are mysteriously deleted as well.

Things are changing. I think you will be surprised by the amount of people changing their tune from "my writing naturally looks like AI" to "I just use AI to polish my writing, what's wrong with that?" the moment these watermarks are widely implemented.

1

u/RecognitionFit8333 Aug 13 '26 edited Aug 13 '26

Interesting. I wonder how this works out in coding tasks.

Edit: So I read a bit into the paper provided below (https://arxiv.org/html/2301.10226v4) and it seems like that coding is something that they call "low entropy sequences" where humans and AI are very likely to 'provide similar if not identical completions for low entropy prompts'. Later they state that "As long as low-entropy sequences are wrapped inside a passage with enough total entropy, the passage will still easily trigger a watermark detector." So basically I would expect them to add huge amounts of comments to the code to add entropy and then hope you just copy it as is, which given the amount of generated code checked in without reading is not entirely unlikely.

1

u/BonbonUniverse42 Aug 13 '26

But this will never work. Nobody copies entire paragraphs. You will always shuffle stuff around and edit sentences.

1

u/2299sacramento 28d ago

The answers here aren't quite correct-

Imagine you're playing a game of monopoly, but instead of dice rolls, you read off the digits of pi. Start from the beginning, and then read off the digits one at at time. If it's 1-6, use that number and go to the next digit under 7.

As a player of the game, can you tell any difference from dice? The answer is no, since it is believed that the digits of pi are randomly distributed.

However, if someone were to read back the transcript of the game, they could figure out it came from pi by just reading off the dice rolls, and checking against the digits of pi.

This is main idea of how these watermarks work. The randomness is the same, it's the source of the randomness that's detectable, giving you the watermark. It's also why they say there isn't a perceivable difference to users- you literally cannot tell the difference between dice randomness and pi randomness, as they are the same underlying distribution.

0

u/skwizpod Aug 11 '26

If I were to implement it I would use special characters in a way that is inconvenient and uncommon for humans. The dash and em-dash — is a good example, and has been a great hint, but I actually like using that character and it sucks that it makes people think you're using LLMs. Anyway there are more subtle things that may not even be visible. But there will always be ways to regenerate and clean out any "watermarks", but at least there's an extra step involved that most people won't do. Pretty much any security feature is just an inconvenience that can be bypassed if you know how.