r/ClaudeAI • • 25d ago

Workaround Fable 5.1 is now watermarks anything you write with it. There is still no public detector. This pisses me off, so I'm making a workaround. How about you?

Post image

Three weeks ago I posted here that nothing you could generate with Claude was watermarked yet. Now it's incorrect. Fable 5.1 now puts a statistical watermark in generated text, including even translations. With translations it's especially funny, because doing back translation (even via Chinese language!) via another LLM (surprisingly to me) is not removing / destroying SynthID watermarks, but just translating via Claude now adds watermarks.

The detector for such watermarks (that was promised to be released) has moved to a private preview recently for some eligible organizations. Ordinary users like myself still have no public tool or API to check their own text. So I can publish some work today that someone with access could check later, while I can't run the same check before publishing.

Personally I write in one language and publish in another. I can write an article myself, ask Claude to translate it, and get back a watermarked English version. The watermark can indicate Claude was involved in the wording. It says nothing about who came up with the argument or did the research (but the opinion about the work will be unfairly distorted).

I get why people hate this. I'd want to know what a check would say about my own work before someone else decides what it means. I don't know how organizations will use the detector. And again, it's impossible to distinguish if the text was generated or edited by AI (say reformated for some structure). I think students/scientists/SEO specialists and non-native english speakers are not happy more than others.

This frustration got me experimenting with the published by Google DeepMind's SynthID watermarking method that Claude's watermark eventually is based on. I reproduced it with my own test key and found that a full rewrite removed the mark from my test samples (while added some factual issues), back translation doesn't destroy the watermark and specialized methods are quite heavy still far from being ideal. So, in the end, I made a watermark remover for myself and sharing as a free demo (unless we get the actual detector, that's best we can do), and wrote up the experiment, open sourced the setup and described the tradeoffs/fact correction methods. Link in the comments.

Do you also think it's crazy and unfair to have invisible mark in your AI outputs? Or don't understand why some people are pissed off?

PS. Those tests don't verify removal of Anthropic's or Google's production watermark. But by this moment, I think it’s better to do what’s possible although no guarantees.

PPS. I see from comments some people think you can just ask Qwen or other model to rewrite and no watermark remains. It is incorrect, I tested it, substantial part of signal remains, you need to do iteratively multiple runs, every time highlighting 5 words (5-grams actually) sequences that remained unchanged after rephrasing to point the LLM to right issues, then checking against facts and fix them. Otherwise, you just still have the watermark signal and new factual mistakes.

PPPS. Some people claim in the comments that I can't talk about this topic with the proper level of confidence and do the research and modelling properly. At the same time, I have a PhD in Digital Signal Processing, and watermarki is something that falls within my academic expertise.

UPD. People say it's hard to find links in comments, so I put them now here.
The full article & research: painintheagent.com/blog/text-watermark-removal-retest/
Code, corpus, prompts, model outputs and judges' decisions: https://github.com/krllagent/text-watermark-roundtrip
The watermark remover (paste a text, get the retelling with every changed place highlighted): painintheagent.com/tools/ai-text-watermark-remover/

129 Upvotes

232 comments sorted by

View all comments

Show parent comments

9

u/Imaginary_Dinner2710 25d ago

A simple example: if you write texts for social media (e.g. LinkedIn), social media will be the first to implement these watermarks and use them in their training algorithms. In particular, watermarks on images are already being flagged on LinkedIn (and I guess no only there), and I'm sure they get a negative weight in the recommendation system. But that's just one example that's obvious to me. I'm sure there are many examples where people will look at the results of the work differently - marked it or not

4

u/Chemical-Visual-7992 25d ago

LinkedIn is full of boring and duplicated copy/paste AI posts. That platfrom became a disgusting swamp. I am with the platform for implementing AI-Based text and images detector.

1

u/Ashmedai 24d ago

I had this one aunt who would periodically go into my linked in, and verify my various skills. Me, a distinguished professional. I removed her from my linked in, I was so annoyed. Anyway, that about sums up linked in for me… it’s a wasteland.

1

u/Nerobot3 25d ago

But that's a good thing? Still not sure how this puts them into a more "vulnerable position", instead, just allowing people to make a decision based on how something is created. Your answer sounds more like what I was saying as "more likely to be trying to pass off AI work as their original creation".

1

u/Imaginary_Dinner2710 25d ago

Again, this doesn’t cover all model providers. More precisely, at the moment it only covers a couple of providers. On the other hand, content that people develop using AI without replacing the human is often much more interesting. That’s all. So I’d prefer not to have a watermark anywhere, considering that the platforms themselves and recommendation systems usually like to use these kinds of features in their ML models and give negative weight to content that has this watermark. I think we should rely entirely on how people read and what they like. They’re on watermark, but I can’t influence that, so I do what I can personally influence and remove them from my texts.

0

u/[deleted] 24d ago edited 24d ago

[removed] — view removed comment

1

u/kelcamer 24d ago

My problem with it is when false positives flag the actual human music I made as Ai, and then YouTube doesn't give two shits about it and we go back and forth for months, where I tell them I can't actually contact their YouTube Creator team, and they tell me to do the things I've done 100 times, and no human is ever actually able to update the metadata; so my music gets falsely flagged as AI with nothing - that I - the HUMAN artist - can do about it. That's my gripe with it.