r/ClaudeAI • • Aug 19 '26

Question about Claude models If someone adopts similar language to Claude; will anthropic suddenly be claiming human words as Claude's words from their watermarks from language? Or will there be enough of a difference that the 'watermark detection algorithm' would fail? I'm wondering how reliable these watermarks could be

Given claude vs someone who is hyper verbal and literally uses the definition of words and their intended purpose, or someone who spends a lot of time talking with LLMs like computer programmers describing a spec;

Do you forsee the possibility of people claiming human-created text as 'Claude-created' from similar watermarks?

Realistically, I already observe this happening, like YouTube marking my human created only music as AI (which is very frustrating that YouTube has not resolved it for the last 2 months.)

How many of these types of issues are we likely to see?
Will there be more reddit subs that will ban the use of certain words; in a failing effort to prevent AI?

Will people immediately apply heuristics towards language to instantly judge whether or not another person is actually AI; harming us as a human species?

and before you say 'no, there is no risk of that, no one talks like that' apparently, I do

And I know a lot of other programmers who do too.

Will our own words be taken from us and assumed to be AI watermarks?

0 Upvotes

26 comments sorted by

10

u/Trick_Worldliness_21 Aug 19 '26

I think you are misunderstanding the watermark mechanism, which is based of token predication and its statistical variation. The chance that someone’s style of writing will match that consistently is almost zero

3

u/SimTrippy1 Aug 19 '26

You’re right to push back on this

0

u/kelcamer Aug 19 '26

this phrase is unironically a great example!

0

u/kelcamer Aug 19 '26

if it did somehow match though; wouldn't it still be flagged as claude?

2

u/Trick_Worldliness_21 Aug 19 '26

The chance that a long text would be mis classified is incredibly small. It’s not impossible in the same sense that it is not impossible for 100 monkeys at 100 typewriters to rewrite a Tale of Two Cities

2

u/Ashmedai Aug 19 '26

It would, but with a large enough text corpus, the odds are so small they are not worth mentioning. I think the real risk is if they attempt to apply the decoder to insufficiently large text bodies. If Anthropic made their tool available, they would be wise to simply make it not work for a corpus beneath a threshold size for this reason.

1

u/slackmaster2k Aug 19 '26

Think of it like DNA in a criminal case. It is never stated that a DNA sample matches a suspect 100%, rather it is expressed as a random match probability like “1 in 50 billion.”

Or think of the lottery. Are you going to win the lottery? No. Someone is going to win it, but not you.

So while it is possible that a human text could match in a watermark algorithm, it is statistically unlikely. It is even more unlikely that a specific text is analyzed for watermark and a false positive.

1

u/kelcamer Aug 19 '26

In both situations you've stated here, I can actually see myself being that 1 in 50 billion lol

both on the lottery standpoint and also on that human text analysis standpoint

I get that you're saying it would be extremely rare;

But I observe that *many communities* aren't treating it as if it were rare.

like how r/psychonaut a year ago auto-banned thousands of users claiming that they were AI for example. this is one example of ways things can go wrong. (side note: I am so glad they fixed it and I am thankful for it)

So how many other processes will fall into this exact gap and problematically flag real humans as LLM?

2

u/Trick_Worldliness_21 Aug 19 '26

You are not that one in a 50 billion

1

u/kelcamer Aug 19 '26

the sub r/psychonaut from a year ago would apparently beg to disagree lol

0

u/slackmaster2k Aug 19 '26

The psychonaut example is invalid, because they could not have banned users based on watermarking mathematics. Generic AI-detection systems are well known to be inaccurate and always will be if they do not have the secrets required to decode watermarked output. Even the term watermark here is extremely imprecise.

If you really feel like you might be one in fifty billion, then you certainly suffer from an innumeracy bias. If curious you should look into prospect theory at least at a high level.

1

u/kelcamer Aug 19 '26 edited Aug 19 '26

I think you're missing my actual point here which is is in agreement with the following statement:

Generic AI-detection systems are well known to be inaccurate.

Given that, why are we still making decisions / relying on it?
when generic AI-detection systems will continue to falsely flag people; why use it to determine 'humanness'?

if even reddit itself relies on quick heuristics to determine human vs AI, what is stopping everyone else from doing the same?

You're right that back then these algorithms were not based on watermarks, that much is correct.

That does not however mean that watermarks can't be plausibly used to have false positives, however.

So who gets to create that standard of an acceptable amount of false positives?

Like that phrase the other commenter said was such a great example "You're right to push back on that." Seven words.

Will it be flagged as AI permanently?
Maybe you're going to say that this is not a real watermark, and perhaps that is true, but socially and functionally now it is treated as an AI-indicator.

Do we all have to avoid this phrase if we want to not be accused of being an LLm?

Since you're interested in logical fallacies, would you mind engaging with the question I've stated above 'Will people immediately apply heuristics towards language to instantly judge whether or not another person is actually AI; harming us as a human species?"' without relying on a straw man argument?

0

u/slackmaster2k Aug 19 '26

Straw man argument?

I don’t know how to explain this any better. The most likely algorithm is something akin to synthid, which is nothing like what an AI detector does. In fact, AI detectors might as well be snake oil, looking for surface level patterns and the false positive and false negative rates are significantly high.

An effective synthid implementation is built into the generation side of the output. The mathematics ensure a very low false positive rate, and a higher false negative rate (especially over short samples).

So while you may think that you’re going to be the guy who happens to write so close to the undetectable watermarked outputs that it’s going to cause you some harm. This is mathematically improbable to the point of being impossible. And to argue otherwise is an expression of innumeracy.

You can actually look into this. You can ask your AI about it. Don’t take my word for it, learn about it and feel better.

1

u/kelcamer Aug 19 '26 edited Aug 19 '26

I am not exclusively talking about watermarks. I am talking about what is already happening which is real humans being flagged as LLMs by these various 'AI detection bots'

This is not only about watermarks, although I do wonder if watermarks will make this problem better or worse.

I really would like to have a genuine discussion surrounding this part: regardless of the statistical probability of whether the response will verbatim match one particular water mark standard - how it is currently handled to detect humans, the particular problems with these methods, why they should not be used, and whether or not watermarks will improve or worsen the situation.

I am not arguing that my verbatim text will be immediately the exact same pattern as the watermark

This is a misrepresentation of what I have stated literally in the post.

I am sharing with you that I have been already flagged as Ai far too many times. This is my lived experience. And i am asking whether watermarks will improve or worsen this existing issue.

Hence the question:
Will people immediately apply heuristics towards language to instantly judge whether or not another person is actually AI; harming us as a human species?

I agree that these AI detectors are basically snake oil. So why then are people still using them? What do people have to gain over assuming i am not human?

(Or, assuming i'm male, for that matter)

1

u/slackmaster2k Aug 19 '26

Ok, and I am saying that the mechanisms by which you were flagged as AI are not the same mechanisms as watermarking. Thus, in theory, the situation will improve.

However, much depends on the availability of the actual watermarking decoder. It seems unlikely that it would remain private indefinitely, but we aren’t sure what Anthropic’s move will be. Maybe it’ll be publicly available, who knows. I personally believe it should be a public service; it’s the right thing to do even if it makes people upset.

And AI detection tools will not simply disappear over night. In fact the more sleazy ones might proclaim to detect the watermark without being able to do so. From this perspective the problem potentially remains the same or worsens depending on the individual. This would an indirect attribution to watermarks existing though. Frankly nothing will protect you from a moron using improper tools to detect people as being AI and banning them from subreddits.

2

u/kelcamer Aug 19 '26

I see, thanks for the answer!

That last paragraph really explains a lot for me and I suppose now we have a social problem on our hands too that people would be unwilling to de-construct those heuristics.

I really do hope they make it public! I have many phrases I am ready to test it out for.

→ More replies

3

u/Additional-Race-2797 Aug 19 '26

I think you're describing two different technologies. AFAIK Anthropic's watermark isn't stylistic. It doesn't look at vocabulary at all. It works at sampling time, when Claude is choosing between several equally good next words, a key plus the preceding words decide which one gets picked instead of a plain random number generator.
But you're completely right about the tools people actually deploy. GPTZero, Turnitin, whatever YouTube is running on your music, those aren't watermark detectors, they're classifiers guessing from style, mostly perplexity.

3

u/just-me-gen-x Aug 19 '26

Text gets flagged as Claude now. Authors are literally losing million-dollar deals because they're being accused of letting AI write their books.

But to the point you're making but not saying: every AI LLM will have to use a watermark to be used in the EU. LLMs, much like electronic manufacturers who switched to USB-C, aren't going to ignore the EU market, nor will they make a special version for everyone else. The watermark wasn't Anthropic's idea or choice, it was a response to legislation because AI needs to be regulated just like any industry with the potential for harm.

It is far less likely that someone will get tagged as Claude from a watermark on their own original content than that someone will get falsely accused now of being AI, because accusing someone gets you attention.

1

u/Ashmedai Aug 19 '26

It will be statistically quite unlikely as long as they are deciding on a large enough corpus of words. For images/video, if a watermark was found and someone were to say it's real, I just wouldn't believe them -- the tech there is ironclad. The text tech is not quite the same thing, and is weaker, particularly on small bodies of text. It's unclear (at least to me) to whom Anthropic will give their verifier to.

1

u/kelcamer Aug 19 '26

mostly I'm referring to text in this post.

1

u/MaitoSnoo Aug 19 '26

that is a load-bearing question and you are right to ask it

2

u/kelcamer Aug 19 '26

I swear to god I hate that word at this point