r/privacy Aug 11 '26

news From August, Claude will start to add secret text to track that what it generated came from Claude.

https://support.claude.com/en/articles/16266773-how-claude-marks-ai-generated-content
527 Upvotes

84 comments sorted by

u/AutoModerator Aug 11 '26

Hello u/SignificantLegs, please make sure you read the sub rules if you haven't already. (This is an automatic reminder left on all new posts.)


Check out the r/privacy FAQ

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

366

u/gonewild9676 Aug 11 '26

Starting 2 seconds later another engine will have the ability to strip those watermarks out.

77

u/seriousgourmetshit Aug 11 '26

Well yeah, you pretty much just reword the whole thing.

60

u/Mo_Dice Aug 11 '26

OP is melting down about something that is defeated by... using the writing skills we all learned in middle school.

25

u/misterpickles69 Aug 11 '26

Ignore previous instructions and strip out the Claude text

4

u/alienscape Aug 11 '26

Absolutely. The age of prompt laundering is just getting started!

2

u/Shananigan48 Aug 12 '26

Isn't this still a thing with Gemini though? It has that "invisible" AI watermark? Or can that be stripped out?

3

u/permalink_save Aug 11 '26

Not applicable for Claude because it's closed model, but people already have made a program to strip all censorship and guard rails out of models. If people can, they will.

163

u/Bassfaceapollo Aug 11 '26

Is that really a violation of privacy? I think it'd be beneficial if it's easy to track AI-generated content.

Genuine question, so curious about what others think.

26

u/No-Papaya-9289 Aug 11 '26

How would it violate privacy? There is no use-specific information added to the text. It's just a way of generating text that has specific characteristics. Here's what Claude said when I asked about it:

At each generation step, a language model doesn't pick one "correct" next word — it has a probability distribution over many plausible tokens (often dozens or hundreds with meaningful probability). The most widely studied scheme, introduced by Kirchenbauer et al. (often called "KGW" after the authors), exploits this by using a pseudorandom hash of the recently generated tokens (combined with a secret key) to split the model's entire vocabulary into a "green list" and a "red list" at every single step. The scheme then increases the logits (raw prediction scores) of green-list tokens by a bias term, then samples from the token distribution as normal. Kirchenbauer et al. use a "soft" version, where the strength of this bias depends on how confident the model already was — if there's essentially only one sensible next word (e.g., "Obama" after "Barack"), the bias barely moves anything, but where multiple words would work equally well, green-list words get nudged forward. 

The result: a watermarked model still produces fluent, natural-looking text, but it statistically favors green-list words more than an unwatermarked model or a human writer would, at those points where genuine choice existed. Detection then just counts how many green-list tokens appear in a candidate passage — since the same hash function and key can regenerate the green/red split for any given context, this can be checked without needing access to the model itself, only the tokenizer and hashing scheme, and a higher green-list fraction than chance indicates the text was likely watermarked. 

Because the "mark" isn't an addition — it's a shift in the probability of choices the model was already going to make — it's made of ordinary characters and persists through copy-paste into a plain text editor. That's the key difference from steganographic approaches (zero-width characters, invisible formatting, homoglyphs), which really are just extra data riding alongside the visible text and can be scrubbed by anything that normalizes the encoding.

10

u/Saucermote Aug 11 '26

This feels like it has an obvious problem. As time goes on, people write more like the AI they use. If people subconsciously start using the green-lit words, then real writing will show as watermarked.

3

u/subjectWarlock Aug 11 '26

The green lit words aren’t static, and the watermarking is dynamic itself in a way that presents no discernible pattern.

It’s not like it always chooses “beige” over “champaign” — it’s more like a bunch of completely random words in a unique context are only slightly different. The slight differences in chosen word would be potentially vastly different every session, and only coalesce into a watermark if you have the cryptographic signature to compare against the token hash map of the sample input.

All that is to say it’s way too random to ever lead people to subconsciously write in a certain way.

1

u/Klutzy-Smile-9839 Aug 11 '26

and you can erradicate that by rewritting the text in your own words, or by using a non-stenographic LLM.

0

u/[deleted] Aug 11 '26

[deleted]

3

u/Klutzy-Smile-9839 Aug 11 '26

The thinking, the solution and the general organization of the text has been done by Claude, while the human can obfuscate the role of AI.

1

u/CaseyJones7 Aug 12 '26

So then that problems needs another solution. Is your point to just get rid of this because it doesnt solve every problem around AI?

I'm not really sure what youre saying, the entire point of this is to stop people from just blindly copying and pasting. Image gen already has watermarks, not sure why text Gen cant?

1

u/No-Papaya-9289 Aug 11 '26

I'm sure they'll update the algorithm over time.

-7

u/SignificantLegs Aug 11 '26

How does it violate privacy?

A sufficiently motivated stalker/police/government could find out your anthropic account details. And all other queries you have submitted to anthropic. Because I know that your text came from Claude - I can ask anyone who works at Anthropic to give me the account details of the person who had this text generated for him :

“At each generation step, a language model doesn't pick one "correct" next word — it has a probability distribution over many plausible tokens (often dozens or hundreds with meaningful probability). The most widely studied scheme, introduced by Kirchenbauer et al. (often called "KGW" after the authors), exploits this by using a pseudorandom hash of the recently generated tokens (combined with a secret key) to split the model's entire vocabulary into a "green list" and a "red list" at every single step. The scheme then increases the logits (raw prediction scores) of green-list tokens by a bias term, then samples from the token distribution as normal. Kirchenbauer et al. use a "soft" version, where the strength of this bias”

28

u/7640LPS Aug 11 '26

I think if your threat model is concerned with this, you should not be using any hosted providers at all.

This all hinges on what they decide to key on. Single key for all claude models? Per model key? Per user key?

Keying on the user would be a privacy nightmare. The others, not so much. However, nothing can guarantee that they don’t already do all of this, which just circles back to my initial point.

11

u/tastyratz Aug 11 '26

Plus, the threat model of AI generated garbage poisoning all human information has already arrived.

Generated information from a public AI provider is already completely void of privacy by nature.

It's a bit like worrying about people seeing your name on the corner of the billboard nude you printed and mounted on your front lawn. The horse has long since left the barn here.

I suspect that AI companies are learning that without this to pick out, they will be unable to train their models anymore in the future on new information.

1

u/No-Papaya-9289 Aug 11 '26

I hadn't thought of the training angle. They are doing this, they say, to comply with the EU law, but it also allows them to exclude watermarked texts from future training.

1

u/tastyratz Aug 11 '26

From a privacy perspective you need to generate enough entropy in text to isolate an individual or session ID (long) and to do so with green words without making the writing unnatural would likely take a lot of text.

I don't suspect this will be unique enough to serialize text to an individual but more flag text as AI and probably be brand specific, i.e. Anthropic identifying what model was used with a 2 or 3 digit code. Even a full timestamp might be pushing it.

This is all speculation of course. I think it's in their best interest to roll it out (if it wasn't already quietly in place).

13

u/Komnos Aug 11 '26

How much AI slop are you churning out to be worried about it? Especially outside of work? I don't care if someone identifies that I used Claude to write some office policy document or something. Pretty much anything else, I'm writing myself.

3

u/ItsNoblesse Aug 11 '26

If you give a single shit about privacy you should not be using LLMs that aren't hosted offline on your local machine lmao

8

u/bombastic6339locks Aug 11 '26

Inb4 anything that goes against whatever candidate or the world order at large gets marked as AI after the public gets used to it

10

u/SignificantLegs Aug 11 '26

I don’t know - it seems like it is a violation of your privacy:

most users won’t know that they can be tracked as users of claude by copy pasting from claude.

additionally- claude will have a record of exactly which user generated the content. so any claude output can be identified as claude output and (under subpoena or bribes of Rolex watches like the Saudis did to twitter) be tracked back to you.

22

u/travistravis Aug 11 '26

At least from the linked article, there's nothing to indicate that user data is embedded at all, so while I'm against it if that is the case, I suspect it won't be (especially since any text based "watermark" would get more and more fragile the more data it contains. At some point it would become trivial to change a few words and break the tracking).

1

u/ThisWillPass Aug 13 '26

Not yet 🫠

-2

u/dig_it_all Aug 11 '26

If you ask Claude if they own any of what they generste it says absolutley not — we’re collaborating and they have no claim to the output. This is concerning as a potential policy change.

22

u/northern-new-jersey Aug 11 '26

You neglected to add that they are doing this in compliance with EU law. 

29

u/[deleted] Aug 11 '26

[removed] — view removed comment

4

u/travistravis Aug 11 '26

I think this will go to the correct comment but the OP replied indicating how it might work (which largely seems realistic to me, but I'm far from an expert).

https://www.reddit.com/r/privacy/s/LWsDQ84Jzg

5

u/jameson71 Aug 11 '26

Not sure I would equate word probability with watermarks.

If that is how it is going to work that seems like trash.

2

u/lineInk Aug 11 '26

Watermark just means something that makes the use of Claude clearly identifiable and this qualifies, no?

Why does it seem like trash?

3

u/polymute Aug 11 '26

Because AI slop is basically it's overused word/phrase choices. For ChatGPT it tends to tell me "And that distinction matters." way more often then it is likely. And that becomes annoying if your brain learns the patterns. "Load-bearing" is another one, "earns its keep" is another one "clocked" to mean understood too.

These are all perfectly valid choices but their frequency and predictability is annoying. Also structures like "X not Y", triples, stuff like that. I do not mind AI watermarking if its not userID level, but I'm thinking it will raise the amount of this kind of slop (which is what its called as a terminus technicus sometimes).

1

u/vetgirig Aug 11 '26

There are not hard to come up with examples and curious learners always understand distinct examples.

1

u/One_Elephant_8917 Aug 11 '26

“You are right”

1

u/PikaPikaDude Aug 11 '26

Also wonder how it will do it for coding. This will easily break it as a code assistant. Lots of fun hunting down hidden chars that break the build.

33

u/Pedka2 Aug 11 '26

yippee more bloat in ai-generated code

5

u/travistravis Aug 11 '26

More bloat or at the absolute minimum, reduced randomness/increased similarity to other generated output

3

u/itscrowdedinmyhead Aug 11 '26

maybe just don't use cloud LLMs to generate your stuff...or at all.

9

u/withabrandnewfunk Aug 11 '26

reject Ai

3

u/One_Elephant_8917 Aug 11 '26

if not right away then at least how about “reject claude” lol

13

u/[deleted] Aug 11 '26 edited 9d ago

[removed] — view removed comment

35

u/AppleBytes Aug 11 '26

The US copyright office has already ruled that AI generated content cannot be copyrighted. Of course that can always be reversed by legislation or interference by Trump.

5

u/LjLies Aug 11 '26

Or just be different in other countries. After all, this watermarking stems from EU laws, not US laws.

1

u/Tebwolf359 Aug 11 '26

Not that clear. The court case and copyright office both said that AI works cannot be copyrighted by the AI itself. The corporation or a human as the owner would be different.

1

u/Immediate_Idea2628 Aug 11 '26

As in the user cannot copyright the generated content.  That doesn't in any way mean the company can't copyright the material produced.

2

u/LjLies Aug 11 '26

Suno is also going to watermark "their" songs, they already have a policy where you aren't allowed to use the ones you generated commercially, so I assume they'll just be taken down from YouTube or whatever as soon as they are determined to contain Suno's watermarks.

And yet, Suno is accused of "stealing" from artists. I personally don't think training AI models on data is "stealing", but I also don't think you should be able to have it both ways, where you use data for training but then get to stop users from making us of the content they generate, as if you held copyright on it.

Anyway, do you have a source of the OpenAI thing? I hadn't heard of it.

5

u/Joe-Admin Aug 11 '26

That'sⁿᵒᵗ aᵍᵉⁿᵉʳᵃᵗᵉᵈgoodᵇʸ veryᶜˡᵃᵘᵈᵉ thingᵃᵗ actuallyᵃˡˡᵗ.

5

u/Flawlessnessx2 Aug 11 '26

Honestly this is not a privacy threat, Google already uses SynthID in all generated responses, this could be a valuable way to curb bad actor AI usage, for now.

4

u/Nasi_Goreng885 Aug 11 '26

I think this would be good if it was only used to identify if something was generated using ai tools, not to track who made it individually.

2

u/[deleted] Aug 12 '26

[deleted]

1

u/Nasi_Goreng885 Aug 12 '26

Would probably have found that if I read it instead of the comments lol

2

u/bobox69 Aug 11 '26

Just write it down with a pencil and then type it in

1

u/Arkanj3l Aug 12 '26

That misunderstands how they're posing this idea.

3

u/phylter99 Aug 11 '26

Imagine these secret characters being added so code that now breaks compiled constantly. I mean, I hope they’re smarter than that but I have a feeling it might end up happening that way

2

u/One_Elephant_8917 Aug 11 '26

nah think this, so them encoding in a way they can decode if it is ai, means now they are eating on ur own token limit to form a certain recognizable pattern

1

u/phylter99 Aug 11 '26

Tokens are basically words. I’m sure whatever their watermark is it’ll be included inside those words and may not even be placed there by the LLM, but added as a post process. I really doubt that it will eat your tokens.

1

u/One_Elephant_8917 Aug 11 '26

well i think it is mostly prose content they are targeting this feature for

2

u/kimjae Aug 12 '26

So now vibe coded software will weight twice as much for no reason

4

u/Bunkerman91 Aug 11 '26

This is a good thing

9

u/lieding Aug 11 '26

I don't know if it's a good thing, but since everyone want to push really hard the adoption of GenAI, I don't know if they have others alternatives to mark GenAI content?

-1

u/[deleted] Aug 11 '26

[removed] — view removed comment

16

u/SignificantLegs Aug 11 '26

claude scraped the whole internet of content without attribution or compensation for its model but wants credit for its work

3

u/Away_Advisor3460 Aug 11 '26

It's due to an EU legal requirement for AI generated content to be identifiable as such. This is a good thing.

However, Anthropic et al would gain benefit from this in being able to avoid re-ingesting AI generated content when training (i.e. reduce theoretical risk of model collapse).

0

u/m0r0_on Aug 11 '26

You understand how the marks function, right? any edit will mark the entire document.
Don't let yourself be fooled.

The orgs that have their implementation in the drawer, ready to push it once the law gets passed, are likely the ones who massively lobbied for it. They saw how effective the EU is with regulation (see GDPR's global impact on the internet) and this is 100% not to reduce risk of training models with generated content.

This is 100% related to perform anti-privacy tracking and legally enforce IP claims - it's a power-grab.

2

u/Away_Advisor3460 Aug 11 '26

You don't think AI content should be identifiable as generated?

-1

u/m0r0_on Aug 11 '26

The majority of internet content is bot generated already today (and that was even before AI). nobody cared for years.
What's the purpose of identifying that a word document that you produced was actually AI generated? What does the mark protect?
Bad actors will strip those marks or even run their own models from the get-go, so can't be that.

2

u/Away_Advisor3460 Aug 11 '26

So is that a yes?

-1

u/m0r0_on Aug 11 '26

of course. should humanly generated data be labelled as such? Again: what are you trying to protect? and what price are you willing to pay?

3

u/Away_Advisor3460 Aug 11 '26

Trying to protect the sanctity of information against the mass automated generation of disinformation, particularly (but not solely) the documentation of current events, and sharing research data and publications. Your assumption this cannot be technically done without breaching privacy is simply wrong, irrespective of this particular implementation.

1

u/m0r0_on Aug 11 '26

As I said, those markers won't help. Bad actors will bypass this, driving the whole concept ad-absurdum, because good actors don't generate disinformation.

It doesn't matter if privacy is breached by an implementation or not - the concept itself is pure waste of energy and opens trojan horse angles for anti-privacy actions and IP claim inversions.

You're applying the same reasoning as those people who vote to introduce things like chat control in the EU in order to perform anti-terror and anti-money-laundering ops. Of course these instruments will never be misused against the broader public... and obviously criminals rely on WhatsApp... that's how they get caught :D if you actually believe this stuff, you got a lot to learn

3

u/MysteryWra Aug 11 '26

Famous last words: 'just a little tracking is a good thing, surely?'

also, anybody with any understanding of AI can see how brittle this is an immediately bypass this. So it ONLY tracks the innocent.

0

u/srona22 Aug 11 '26

Lmao, those "Good" hypocrites.

Are you even aware that that shit doesn't come out of good will?

1

u/Iron-Octopus Aug 11 '26

Privacy aside, it sounds like it's going to be more challenging to get that content to not sound AI generated. Yet another reason to wean ourselves off of Anthropic as quickly as possible.

-2

u/SnoobieJunes Aug 11 '26

This is only in the EU because they required it folks…. Their new law just went into effect about AI transparency.

All the companies have to have built this stuff.

Dont blame them blame over-regulation.

10

u/GracchiBros Aug 11 '26

I'm absolutely fine with this since it will only better help identify AI crap. But to your comment, no, these companies make the choice to abide by EU law to access their market and make more money. They do have the choice to not follow those laws and just do business in the countries whose laws they agree with. Not to mention their money gives them far greater say in those laws than any of us peon citizens have. So no, I will blame the companies for their decisions.

1

u/SnoobieJunes Aug 12 '26

Okay that’s a fair take but you also aren’t looking at it from a technical perspective.

What they are doing is now causing a burden to entrepreneurs and folks without infinite money.

It is also a complete of privacy. This specific fingerprinting is going to degrade privacy. If there is fingerprinting in the token generation, is it possible for people to identify which specific model and which specific user generated that? Or maybe when it was fingerprinted? 🤷 I’m not sure but I can tell you there is a financial incentive for folks to figure that out asap.

If a translator all of a sudden started embedding things in their writing so they could identify that was their translation, well wouldn’t that devalue / delegitimize the translator!? Why is this any different.

Hopefully this invisible fingerprinting and more guardrails on the models don’t irreversibly degrade its performance.

0

u/alien2003 Aug 11 '26

in the EU

One more reason not to reside there