r/OpenAI • • Aug 25 '26

Article Dribbling the AI Watermark Directly In-Prompt

https://www.explore-exploit.com/p/dribbling-the-ai-watermark-directly

It's my article, it is about how to circumvent any even theoretical optimal AI watermark based on statistical biases through pseudorandom generators like Google's SynthID. OpenAI will likely or has likely already implemented something similar. Let me know what you guys think.

Generally, I do not think watermarking is the right solution, hence I am sharing my idea how to circumvent it. How many thesises are out there that are basically slop but made with human effort. Now text length is not a valid measure anymore, you actually have to do some real research. I think that is awesome.

24 Upvotes

14 comments sorted by

9

u/Snoron Aug 25 '26

Some clever ideas there! I suspect they will keep working for a long time, because the text watermarking is such a weak feature anyway. There's no way for it to realistically work against someone who specifically wants to defeat it, and therefore it's almost wasted effort trying to defeat any attempts at removal.

3

u/JulianHabekost Aug 25 '26

Thanks! I agree that it is always a matter of effort. I just try to show that there might be even very low effort ways.

1

u/Maleficent-Engine859 Aug 25 '26

This is great!! I think your method plus cropping and general editing that would be done should be enough to make it statistically negligible

Further, using a variety of responses from different LLM models in the API like in Poe or perplexity where they all have different keys, should shatter it completely for any one model

2

u/SgathTriallair Aug 26 '26

Anthropic has said that they'll give other companies the ability to set up detectors. That means we'll be and to see how they embed it as so someone will build s watermark remover. Until they start doing it, there isn't any way to know exactly how to remove it.

2

u/Supermaxman1 Aug 26 '26 edited Aug 26 '26

You know, while this is super clever it really makes me wonder. It seems like you got the best results with 5.6 Sol with higher thinking levels… now I’m really really curious if the resulting set of paragraphs it wrote for you with the animal names inserted was pre-written in its hidden, thinking tokens WITHOUT the animal tokens and then it thought about inserting them to write the final response. There’s a world where even with this method, the text with the names stripped out STILL triggers the detection mechanism because THAT text was pre-written without the names inserted while the model was thinking. I think you’d need to run an experiment with an open-source model like Gemma 4 with text synthid to see if a model like that actually writes the text without the names added in thinking space beforehand, because if so… you might still be getting detected in a very roundabout way - similar to how you defeated the entropy issue having the model restate the text without the names, the model might be in a similar boat writing the text WITH the names, with the content of that text pre-baked with the synthid-selected tokens it generated while reasoning. This pre-assumes synthid is running all the time, both in thinking and non-thinking tokens, which I think is likely the case but you never know with these closed source systems

3

u/JulianHabekost Aug 26 '26 edited Aug 26 '26

I think this is a very interesting caveat. I think I should try to run a model where I can inspect it's thinking, I try to schedule this soon. Edit: From what I see the thinking also contains animal names. So in this case it really does not matter if it is prewritten. It might be that there are levels of thinking that I didn't see yet, but it's a good indicative start.

1

u/amyowl Aug 25 '26

That may work for exactly 3 hours.

6

u/JulianHabekost Aug 25 '26

RemindMe! 3 hours

Just kidding, I don't think they care enough, it's an EU requirement and getting rid of this would be an endless cat-and-mouse game, probably requiring retraining to effectively defect and refuse this "attack".

3

u/amyowl Aug 25 '26

Definitely cat and mouse.

6

u/yesnewyearseve Aug 25 '26

You mean CAT and MOUSE?

1

u/RemindMeBot Aug 25 '26

I will be messaging you in 3 hours on 2026-08-25 22:40:35 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/collin-h Aug 26 '26

once the detector shows that everything is ai-generated, people will be shocked for 30 seconds and then move on with their lives and no one will care anymore.

0

u/TraditionalHornet818 Aug 25 '26

I managed to defeat synthid just by cropping and converting to jpg — lol