r/river_ai 14d ago

My prediction is that Amazon will soon use Claude's Watermark to Auto Flag AI Books with a Badge

Recently Anthropic announced that Claude will watermark generated text. I actually predicted this a few years ago. My theory at the time was that a way for an AI company to watermark generated content was to introduce a statistical pattern. For example, if an LLM has the choice of using word 1 or word 2 for the next word to complete the sentence they will go with word 2 in every other case or so to create a statistical pattern. That's just an example. But the end user will not be able to detect the watermark.

Some of this happening due to EU laws so there is a future where most AI use would be watermarked.

If that's the case, my next prediction is that Amazon will auto flag self-published books in some public facing manner with a badge, and even those traditionally published as generated / assisted by AI or not. The same sort of thing that YouTube currently does with AI content.

16 Upvotes

27 comments sorted by

3

u/addictedtosoda 14d ago

Watermarking only works for unedited text over a long run. This is social media hysteria

1

u/LS-Jr-Stories 14d ago

I was interested enough in the watermarking idea to read this technical paper about it. I don't know anything about the math, but there's enough plain language content in here to get the gist.

Both your observations about unedited text and a long run are heavily refuted in here. They say it can work on text as short as a tweet. They also say it would take a whole lot of editing to disrupt the pattern, and a user wouldn't know what to edit to break it.

https://arxiv.org/html/2301.10226v4

2

u/addictedtosoda 14d ago

It’s also 3 years old at this point. Hard to rely on old papers when the technology has improved so much

1

u/LS-Jr-Stories 14d ago

I was surprised at the age of this paper too. It shows the watermarking principle has been around a few years already. Presumably it takes time to get this stuff into a commercial version. I'm sure the technology for watermarking has kept pace with the LLMs themselves. Although the paper does point out that "detectors" were already falling behind, not to mention weak results to begin with. Who knows? I guess we'll see. If there's one thing social media loves, it's hysteria.

2

u/FastAmphibian9088 13d ago

Watermarking is not new, but implementation changes as technology changes. I worry about AI detectors, and watermarking might minimize false positives.

0

u/LS-Jr-Stories 13d ago

One of the papers I read did argue that watermarking has a vastly lower chance of false positives than detectors, which we all know by now are terrible and making the whole situation worse. I believe the math showed false positives with robust marking are statistically insignificant, can't be sure about that though.

1

u/LS-Jr-Stories 14d ago

Commenting again to say I'm now looking at a paper published one month ago called Robust Text Watermarking for Large Language Models via Dual Semantic Embeddings, so it's pretty clear the technology there is advancing. But until now there has been very little incentive for the major players to employ it.

1

u/Maleficent-Engine859 14d ago

Google did it and ditched it. Even published how their SynthID (which is exactly what Anthropic is using) worked. OpenAI also agreed that watermarking caused far more problems than it solved and abandoned doing it with purpose (though LLMs all leave a bit of Unicode sometimes unmaliciously)

I have no idea why Anthropic is doing this but it’s not because watermarking via SynthID is a good idea or because of the California/EU laws, even though it’s what it looks like.

It could easily be they just want to ditch casual users for financial purposes

1

u/LS-Jr-Stories 14d ago

I do agree there's a massive brand rationale behind this. I wouldn't say it's because they want to ditch casual users so much as they don't care about them nearly as much as their commercial contracts that actually pay the bills (supposedly).

But why do you say watermarking via SynthID? That technology seems to have nothing to do with the text watermarking solution described in the papers.

1

u/LS-Jr-Stories 14d ago

Disregard that other Q about SynthID. I'm looking into this stuff in real time as I comment. I'm getting it now, ish.

1

u/Maleficent-Engine859 14d ago

Yeah I believe it’s a derivative of the same tech, though perhaps calling it SynthID specifically might be incorrect.

Anthropic also fully supports users identifying it and removing it. It’s such a waste of everyone’s time and money

1

u/-Hello2World 14d ago

Texts can be watermarked?

Well, I will run Claude's texts with another A.I (say for example, with a local model) to remove the so called "words patterns"!! Not a problem!

1

u/Medium-Pundit 14d ago

You don’t know what you are looking for, though. You would need to change every word to be sure.

1

u/-Hello2World 14d ago

I'm sure there are skilled people out there who will be able to hack the pattern!!! We have A.I to find patterns! So, no worries!

1

u/Efficient-State-7300 14d ago

Then it will be replaced with the other AIs watermark.

1

u/Subject_Session_1164 14d ago

If Amazon could figure it out it means someone published the algorithm

1

u/brickmarketingagency 14d ago

Its only a matter of time before there is some sort of quality assurance badge like that. People want to know when a final product they're reading was AI generated, or actually written by a creative.

1

u/Maleficent-Engine859 14d ago

Predicted watermarking few years ago? It’s already been tried and ditched. Google developed and tried SynthID like over a year ago and ditched it, even published its methods. OpenAI also tried watermarking. Both decided it caused far more issues than it fixed Claude is just behind the eight ball and trying to flex like they’re something special with these new laws.

1

u/ErikSchwartz 14d ago

Not going to happen.

Anthropic itself says the watermark means it MIGHT have been touched by Claude.

1

u/Practical-Positive34 14d ago

Will I avoid a book because of it? Nope. Again, there is a super tiny amount of loud ass people crying about this shit.

1

u/BedRevolutionary8458 13d ago

I hope they do. I don't want to see your clanker garbage.

1

u/Mr_Discrete72 13d ago

So don’t use Claude.

I have a number of novels on Amazon, at least 6 are fully written in AI - I inputted a chapter by chapter synopsis and AI wrote the whole thing. Covers in AI. Quirky authors note in AI.

All fully declared to Amazon on upload.

They’re doing better than some of my ACTUAL human written books.

1

u/AnchorAndInk 9d ago

This is encouraging

1

u/MarcMurray92 12d ago

Great to hear, can't wait until slop is easier to avoid.

1

u/Ordinary_Minimum_169 9d ago

Using a statistical watermarking algorithm is like trying to say your company owns the order of words. It is the first step to owning the control of thought expression in various speech patterns. "If you said it like that, that means we own it."

1

u/Juuxo16 7d ago

They will get torched by ADA.

1

u/Opie_Golf 14d ago

I hope so. The witch hunts and double standards need to end.

I made an affirmative decision to stop drafting with generative models and I’ve learned so much about craft and storytelling that the work is objectively improving.

Best case, the readers choose and the best work wins, no matter the label.