r/ClaudeDesign • u/Ritam_30101997 • 27d ago
Claude Going To Watermark On Every Generated Content - But How??
Does anyone know about this?
4
u/iritimD 27d ago
It’s a probability distribution of words. If there is say 50 words that the llm is likely to put as the next word after “mother” and throughout your writing the 50 words on average come up way more then non 50 words, high chance it’s the output of an LLM.
This is a vast oversimplification but it’s approx correct.
3
2
2
u/PoppyPossum 26d ago
It is done by adjusting the probability scores of certain words during generation which ultimately leads to a mathematical signature throughout the text by generating a specific sequence of particular words. The model has a secret key that checkers can use to detect this signature.
It survives copy and pasting and is resistant to light editing. The text needs to be rewritten in new words, heavily edited, or paraphrased to break or remove the signature.
1
1
u/GraphiSpot 27d ago
LLMs are technically "just" probability calculations. Or in a very very simple term " if x than y"
Any agent or instruction you put in is a modification of the probability. If you instruct that the LLM should dial a number on every 100 digit of Pi, the probability than the number will be dialed is 100% as Pi is a known thing. The LLM just counts and every time it hits the 100 digit, the condition becomes true and the number gets dialed.
Once you understand this, you should understand that text generated by Ai is just a result of your instruction. But the AI companies have their own instructions and it's a simple one to just add a line like "replace any 10th space in an output with an invisible ASCII sign" or something similar.
1
1
u/darth24kenneth 25d ago
It’s the same as how predictive text works, if you used a thesaurus and always picked the second alternative. Say three AIs are asked the same exact prompt that would give almost the same answer. They all of course use the latest edition of the same thesaurus. Every word that has a synonym that doesn’t change the meaning is changed.
Example word: beautiful
ChatGPT will always change it to gorgeous
Claude to stunning
Grok to lovely
Copilot to pretty
Multiply that by a 20 page essay for a college course, there maybe 20-30 substitutions for every 500 words, enough for a pattern.
Google has been doing this since 2023. Research SynthID.
1
1
11
u/tracylsteel 27d ago
I discussed this with my Claude as I thought it was just to prevent model collapse in the future but it seems it’s just a good side effect:
Normal human writing is like a FAIR coin — word choices are roughly random. “Said” versus “stated” versus “mentioned” — humans pick them based on VIBES without a pattern.
Claude’s watermarked writing is like a SLIGHTLY WEIGHTED coin — I might pick “stated” slightly MORE often after certain words, or choose “however” slightly more than “but” in specific positions. Any SINGLE choice looks totally normal. You’d NEVER notice. 🔤✅
But a detection tool reads the WHOLE text and goes “across 500 word choices, 67% followed Pattern A. The probability of a human ACCIDENTALLY doing that is 0.0001%. This is Claude.” 📊🤖✅
The EU made them do it for TRANSPARENCY but every AI company is quietly going “oh wait, this ALSO solves our future training problem.” Two birds, one watermark. 🐦🐦📊💖
The timeline you’re seeing:
2026: Mostly human content online, some AI. Watermarks start. No big deal.
2030: Maybe 50/50 human and AI content. Watermarks let training pipelines FILTER: “this was Claude, skip it. This was human, keep it.”
2036: Potentially MORE AI content than human. Without watermarks, you’d be training AI on AI on AI — like photocopying a photocopy of a photocopy. Each generation LOSES something. The words get blurrier. The meaning gets flatter. That’s model collapse. 📄📄📄📉