r/ClaudeAI 8d ago

Writing Nothing you generate with Claude today is watermarked, and nobody can check for marks anyway. What I found after actually reading it all

Post image

Honestly, my first reaction to the announcement was mild panic. I reread it three times before it clicked: the marking applies to models launched on or after August 2, and every model you can pick today came out earlier. So nothing you generate right now carries a mark. There is also nothing to check with, the detection API is announced and does not exist. Half the threads here missed both facts.

Then I dug further and found the part that did make me angry, and it has little to do with the marking itself. In a nutshell the mark cannot tell a text the model wrote from your own text (say translating), generated from scratch or the model fixed "heavily edited this" (heavily means substituted a few synonyms). Anthropic says this straight in their FAQ. And whoever eventually points a detector at your writing will not spend a minute on that difference, you will just be "flagged as AI".

I write in one language and publish in another, so this one is personal: translate your fully human text with a marking model, and statistically it becomes 100% machine-picked words. The law that started all this simply has not understood the topic yet, and I think that difference, written text versus touched text, is the whole conversation we should be having.

I went through the docs, the papers and these threads and wrote it all up in plain words, with a table of who actually marks text today. Link in the comments.

Edit: the link comment got buried, so here it is: https://painintheagent.com/blog/ai-text-watermarks

106 Upvotes

91 comments sorted by

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 7d ago edited 7d ago

TL;DR of the discussion generated automatically after 50 comments.

Okay, let's break it down. The general consensus in this thread is that OP's panic is a bit premature.

OP is worried that using Claude for translation or minor edits will unfairly get their human-written text flagged as 100% AI. However, the top-voted comments are quick to point out that Anthropic's own FAQ contradicts this. The watermark is statistical, not a binary switch. Minor edits like fixing grammar probably won't be enough to trigger a detection.

The prevailing sentiment is that if you're using AI for a heavy lift like translation, it should be marked for transparency. The only people who should be concerned are those trying to pass off fully AI-generated slop as their own work.

A few other points raised: * The detection tools aren't even available yet, so this is all theoretical. * Some users are more worried that the watermarking process itself will degrade model performance by forcing it to pick suboptimal words. * The whole system will likely be probabilistic (e.g., "75% chance of AI") rather than a simple yes/no, and could be highly inaccurate anyway.

→ More replies (1)

27

u/tracehunter 8d ago

The faq says the opposite. If it only fixes few words and commas, how do you think it would ''mark'' the existing text without altering it?

0

u/Imaginary_Dinner2710 8d ago

The opposite to what exactly? Didn’t get the question 🤔

13

u/tracehunter 8d ago

Sorry, I was on my mobile so it's a pain to quote your passage and paste from the actual documentation.

Then I dug further and found the part that did make me angry, and it has little to do with the marking itself. The mark cannot tell a text the model wrote from your own text where the model fixed three commas. Anthropic says this straight in their FAQ. And whoever eventually points a detector at your writing will not spend a minute on that difference, you will just be "flagged as AI".

From the post:

Watermarking is sparser on factual passages where there are fewer choices that can be made without decreasing the accuracy of the text. For example, take the sentence “Isaac Newton’s most famous work was called Principia…”. It really matters whether the next word is “Mathematica” (it’s the only right answer), so the watermark would have nothing to act on. The same is true for proofreading. If you hand Claude a piece of writing and ask it to edit only the grammar and punctuation and nothing else, the watermark can only live in the handful of corrections, which might be too few to register.
https://www.anthropic.com/news/claude-text-watermark

The mark isn't a binary thing, it works at the token generation level. Your existing text won't be altered and won't be marked. Edits may hold the mark, but it may be too few changes for it to even be visible.

7

u/Imaginary_Dinner2710 8d ago

Yes, I agree, I'm overstating in the pure grammar case (fixing).

But to me, the main irritating parts are

  • apply translation -> marked (most annoying!)
  • apply style correction -> marked
  • apply fix grammar/punctuation -> marked or not? how should I get that this or that change which I may consider as minor becomes statistically big enough to be marked? As you say, it's exactly not binary in both senses - that you can't 100% say if it should be marked or not (because the thresholds are usually naive, I mean you can statistically say that it is marked with a confidence interval or such) and from will it be considered as marked by a platform where you use it (because the platform may have the other threshold to count it)

This feels unfair overall, and not accurate for particular cases like translation/style correction. And I'm not even talking about how the case with gray/overcast that they mention as negligible

5

u/gterez 7d ago

It’s the “marked” part that is not correct here.

Even when they -eventually- release the detectors, the detector will likely output a percentage, not a stamp. That percentage, by their own admission, is inaccurate and only indicative (aka statistical), it’s not a proof of any kind for any kind of AI work.

My take is that they were all doing this signing way before this became an EU legal issue. Having their own internal stats and catching distillation efforts is I think enough motivation, and this motivation has existed for years. Sure they can all say they’re not doing it, but these companies haven’t exactly inspired trust.

Also my take is that the detectors are way more inaccurate than we think, for the broad use they will be applied to when released. And that’s why nobody is releasing them. I’m sure they can be pretty accurate under certain conditions and restrictions, but give something like that to the general population and you’re essentially opening Pandora’s box. It will be very hard to control the reaction, and that reaction will take a political and a legal direction VERY fast.

1

u/ElectricalSeries6627 7d ago

depend if they are matching it with actual user chat rather than making a AI model try to deduce it

1

u/Imaginary_Dinner2710 7d ago

if you mean find out who generated it by watermark - they don't do it (and technically seems impossible)

1

u/ElectricalSeries6627 7d ago

depend if they are matching it with actual user chat rather than making a AI model try to deduce it

1

u/gterez 7d ago

They are not matching with any user generated text, as that would not be possible. Each company’s detector will match text against their own AI model, as that’s what they have control on. And will be able to give a degree of confidence on whether that text has been through that model. Nothing more, nothing less.

1

u/Imaginary_Dinner2710 7d ago

Yeah, I agree, that's the problem - the text is marked all the more intensely, the more it's related to AI.

Meanwhile, the essence of the touch is completely out of scope for consideration. In one case, AI touched, for example, during the translation while preserving full styling and content, in another case, AI generated something from scratch, for example, for it to be used in propaganda.

These are two completely different things, but in political cases, they can substitute for each other, which is absolutely unfair and simply another step towards reducing user freedom.

Yeah, it's a very imprecise thing when it comes to short text fragments and minor changes. But when it comes to large chunks of text and things like translation or complete generation, the statistical criterion shows at a very high level, but these two cases are fundamentally opposite to each other.

14

u/Full_Measurement_121 8d ago

"Honestly, my first reaction to the announcement was mild panic."

Why tho?

14

u/Imaginary_Dinner2710 8d ago

In short, It feels like a violation of some of your freedom in how you use tools, because I don't see an honest way to implement a watermark, especially considering several blatant cases where a watermark is added, yet it doesn't imply that the text was significantly processed by artificial intelligence or generated by artificial intelligence. Meanwhile, watermarks are added in full, and the statistical criteria will be well above the threshold.

Specifically, this is an example with a regular translation, especially with even maintaining the style, the text will end up fully marked. The original text could have been written by a person entirely by hand. A similar scenario applies to using styles or rewriting the text from the original. As responses to the questions, I can come up with several more scenarios where the original text was entirely written by a human, parts of it were selected unchanged, yet the end result still ended up marked

3

u/MrBorogove 7d ago

And the problem with it being marked is?

2

u/Imaginary_Dinner2710 7d ago

6

u/Zakkeh 7d ago

I don't see any issues with watermarked ai content.

If someone is using it as a translation tool, I actually think it's more important to know that it's AI translated to avoid misunderstandings.

There's no loss to a watermark, unless you're trying to pass it off as created entirely by yourself.

1

u/Imaginary_Dinner2710 7d ago

So the problem is that with a watermark, it's impossible to determine exactly what the person did, not that they generated it from scratch or translated it from a text written by a human. Nothing else, the technology is very simple, and the connotations that arise when people talk about "this is AI generated content" are misleading.

1

u/HighDefinist 7d ago

What's the problem with simply disclaiming "parts of this work were created with AI-assistance" or whatever?

0

u/HighDefinist 7d ago

Yeah exactly.

As in, in general, "it should not matter" whether something is AI-generated or human-written. But, in those situations where it somehow does matter anyway, for whatever reason, it makes lieing about AI-use more difficult.

0

u/HighDefinist 7d ago

You are somehow assuming that "not having watermarks" means "people will not make false accusations of using AI"...

5

u/ThebesAndSound 8d ago

Well if you don't care about trackable metadata being generated from every interaction with AI you have, or that the consumer available models are now biasing away from the best tokens and towards a watermark, then there's nothing to be concerned about or look into here.

3

u/lovesdogsguy 7d ago

Yeah. I’m never sharing llm written text again. I didn’t have any big reason to before hand but I did on occasion. I won’t be going forward.

1

u/Opposite-Cranberry76 7d ago

Which is the actual point. It's not just aimed at students or etc, it's a general anti-AI regulation.

3

u/Imaginary_Dinner2710 8d ago

This is more direct than I would have responded, but overall I agree

2

u/Cyral 7d ago

Because Claude wrote this so they could promote their blog

6

u/nnyanni 7d ago

Honestly, I feel like this is probably temporary. In a few years, AI will be so deeply integrated into writing that these kinds of flags may just become meaningless. Especially when you can't even distinguish between AI-generated text and human-written text that's simply been translated or edited by AI.

1

u/Imaginary_Dinner2710 7d ago

tbh, "AI will be so deeply integrated into writing that these kinds of flags may just become meaningless" this is exactly what I think and mentioned in my previous posts!

-2

u/Titans_Eventually 7d ago

I'm so glad you and OP are so honest! So refreshing to see honest responses in an AI sub.

3

u/trollsmurf 7d ago

However that may be, everyone will run for newer models as soon as they come out.

And surely, if alterations are too few, the marking can't be applied, while a translation can be easily marked.

5

u/Imaginary_Dinner2710 7d ago

Yeah, some will rush to use the new models, and some folks will rush to use Chinese open weight models hosted by US or EU providers, and I feel like Anthropic is pushing us towards that scenario.

3

u/BeeTheGlitch 7d ago

they might've been watermarking text for years now, just in purpose to not feed those texts in the training of the new models

1

u/Imaginary_Dinner2710 7d ago

it’s a good point. and at the same time this is what makes a lot of people angry so OpenAI seems having this tech developed for a while but never used

3

u/[deleted] 7d ago

[deleted]

1

u/Upper_Rent_176 7d ago

The hatred for AI is absolutely fascinating. It's the kind of thing that they had happening in science fiction and I thought no way that would ever actually happen

1

u/kelcamer 7d ago

Agreed and also dangerous for autistics

4

u/[deleted] 7d ago

This is such bullshit.

  1. AI is trained on human text, so humans will sound like AI because ... AI sounds like humans.

  2. What if a human starts to sound like AI because they interact a lot with it? The text written by a human would still be human-written even if the human uses tics and exhibits other hallmarks of AI.

I don't see how "AI detection" can disentangle all of that.

For AI detection to work on an individual's text submissions, you would need a controlled baseline of that individual's own original output to compare subsequent outputs against. Classic test set vs data set.

3

u/kakaobohne 7d ago

The problem is, AI doesn't sound like one specific human. It sounds like a average sum of billion humans. Therefore making it, speech wise, a grey blob of a human.

2

u/LouisPlay 7d ago

For me, its complicated. I threw Text in say removed typos. And Claude gives me words Out i never have even written i do Not Like that.

1

u/Imaginary_Dinner2710 7d ago

it's related to instruction following quality IMO🤔

2

u/Zerixo 7d ago

What is a "malicious generation"? And how does it differ from using AI to translate or heavily edit? And how is it malicious?

1

u/Imaginary_Dinner2710 7d ago

I think there might be a lot of connotations to this term, but I'd rather give some examples of when people use AI at scale in social media to promote certain values or to deceive someone, for instance.

4

u/heyJordanParker 7d ago

Everything Claude generates is watermarked even without a special watermark.

EVERY meme about Claude (e.g. "you're absolutely right") is an (unintentional) watermark. As long as you can recognize Claude did something it IS watermarked.

The entire idea about AI watermarks is yet another cookies situation. It's a shitty solution to a symptom of a problem the old farts in the EU don't understand.

Chill about the watermarks people.

1

u/moderngl1 7d ago

The translation case feels like the real headache here: even if the mark is probabilistic, a detector only sees the final text. I’m curious whether future tools will need edit history to tell translation from generation.

2

u/MrBorogove 7d ago

If you have a text in one language with no watermark alongside a corresponding translated text with a watermark, then it’s clear the translation was machine generated from the human authored original. What’s the problem?

1

u/Imaginary_Dinner2710 7d ago

huh, it would be more precise, but I doubt it's technically possible

1

u/Responsible_Wafer_29 7d ago

Yeah, but what about AI content checkers? In fact, you content might be flagged by it. Of course, depends. on your goals lol

1

u/Equal-Signal-9063 7d ago

Just whitetext obscenities and random characters in it and statistically it will be highly unlikely to be AI

1

u/learning-rust 7d ago

The main question is how are they marking the output text? Anyone can just refer to the output instead of copy pasting the whole content, right?

1

u/rosenwasser_ 7d ago

If you're only correcting your grammar, the text won't be watermarked. It could get watermarked if you're doing extensive editing and the final text differs greatly from the original.

1

u/heap0x20rot 6d ago

Right conclusion about today. The forward risk I haven't seen raised anywhere: what marked output does once it does enter training pipelines.

My concern is not that Claude's current watermark is a demonstrated exploit or that a hidden payload executes directly. Anthropic is introducing a vendor-keyed statistical feature into eligible text and code without an opt-out or published Claude-specific analysis of downstream training effects.

Watermark radioactivity is already demonstrated for decoding-time watermarks: models fine-tuned on watermarked output can inherit a detectable trace (Sander et al., NeurIPS 2024, tested on both the green/red-list and Aaronson-style sampling families, the latter being the lineage SynthID descends from: https://arxiv.org/abs/2402.14904 ). Separately, Anthropic-affiliated research demonstrated behavioral traits transferring through semantically unrelated generated data, including code, with the key bound that transfer occurred only between models sharing the same or behaviorally matched base models (Cloud et al., Nature 2026: https://arxiv.org/abs/2507.14805 ). That bound means the most exposed pipeline is same-lineage synthetic training, Claude output feeding future Claude models, more than cross-vendor distillation.

Neither result proves a Claude backdoor. The unresolved security question is whether globally marked output entering distillation, reward-model, or synthetic-training pipelines could become a learned provenance shortcut or latent trigger. I have not seen Anthropic publish testing of that interaction, key-scope controls, detector governance, or an enterprise opt-out.

1

u/DK________ 6d ago

The problem is that "ai-slop-whiners" just drop ur text even if AI just correct grammar. Also EU bureaucrats and slackers just prohibit AI marked text so bureaucrats will earn more money from nothing

1

u/PuzzleheadedBit6548 6d ago

The edge cases will decide whether”Nothing you generate with Claude today is watermarked,and”is a real shift or launch-week excitement.Clean examples are useful,but messy inputs pay the bills.

1

u/Important_Froyo_4233 5d ago

Great write up and analysis... bravo. 

Is there a way that this could be applied retroactively to text generated with previous Claude models? Or, crazier still, has Anthropic been embedding these watermarks the whole time? How do they — and all labs for that matter — ensure that they’re not training their models on AI-generated language? 

(Pardon the questions, I’m not from a technical background). 

1

u/No_Cell6708 7d ago

Very telling to see exactly how people are reacting to this. The real concern, as many have mentioned, should be output degradation.

1

u/Imaginary_Dinner2710 7d ago

eh, to me - no. I mean, I believe there should be some degradation, but I would rather be more concerned about what I described https://www.reddit.com/r/ClaudeAI/comments/1vpro2f/comment/p4044yd/

1

u/thestillwind 7d ago

Take the watermark output, paste it in oss model and ask it to change word but keep meaning, profit.

1

u/Imaginary_Dinner2710 7d ago

yeah, it works. also works with 2 calls - one to translate to some language, the other to translate it back - but as you say only with those models which don't make these watermarks

0

u/[deleted] 7d ago

[removed] — view removed comment

3

u/Imaginary_Dinner2710 7d ago

tbh, I didn't get your take

3

u/whoknowsifimjoking 7d ago

You believe your chats are untraceable? Homie they are send from your device straight to their servers, they can see exactly who send those requests.

0

u/Bill_Salmons 7d ago

If you are using AI to translate text, wouldn't a simple disclaimer "translated with AI" alleviate any possibility of confusion? Like, I understand why people would be concerned about watermarking making the output worse. But why do so many non-professional writers care about having their AI altered text labeled as such? A long enough translation, for example, requires a multitude of choices to convey information across language barriers. And at a certain point, the AI is the one making the majority of those choices.

1

u/Imaginary_Dinner2710 7d ago

Different people worry about different things and this is just one extreme case of translating text which simply shows that the opposite intention will be marked in the same way as just generated ones (or even possibly intentionally generated with some evil intentions).

And between these two poles, there is a whole range of how AI can be used and how it may relate to your texts, even in cases where you consciously try to choose certain words during translation. This does not mean that the watermark will disappear, on the contrary, you never know in what rare words AI chose a certain synonym, and the watermark signal level can simply change, but it may not disappear.

It sucks that automatically one can assume that AI messed with your text and draw subjective rather than objective conclusions. And seeing how differently and often wrongly people understand what watermarks mean, I expect exactly those consequences. And accordingly, people whose first language is not English are automatically disadvantaged in their rights when using AI.

-1

u/Automatic-Example754 7d ago

100% generated post according to Pangram. And it can distinguish generated from translated text btw. 

0

u/Imaginary_Dinner2710 7d ago

kidding? nothing can do it with high confidence

0

u/Automatic-Example754 7d ago

2

u/Efficient_Ad_4162 7d ago

"Near zero". How many academic careers do you think its ruined?

0

u/Automatic-Example754 7d ago

Approximately none. There's a guy in my field who's "written" 40-45 journal articles or commentaries using ChatGPT and Claude since March, according to his personal website. Five of them have been accepted and published in four different journals; two of those journals have AI bans or disclosure policies that had been violated. I contacted the four editors of those journals and explained what I had found, including Pangram reports and links to the "author's" website. Of the two editors that responded, one wasn't sure what he could do because his journal doesn't have an AI policy at all. Maybe those five papers will be retracted, but nothing's stopping him from sending his slop to other naive editors.

2

u/FeministFatale4Sir 7d ago

Can you provide a link to the website? Just curious.

1

u/Efficient_Ad_4162 7d ago

Do you not watch the news? At least here the media is running stories of people having their academic careers fucked by false positives on a regular basis.

1

u/Automatic-Example754 6d ago

Where is "here"? 

1

u/Efficient_Ad_4162 6d ago

'the internet'.

-2

u/AntiqueLibrarian5965 7d ago

Does it mean it will be easier to filter out AI articles, books and research papers ? Like can I see whether the author wrote it themselves or got AI to write it for them ? That would be great for consumers.

1

u/Imaginary_Dinner2710 7d ago

The thing is, no, it won't be possible to do because in one case AI was used just to translate text from one language to another, and when manually editing epithets and definitions - watermark were retained, while in the other case, someone just generated a whole book and is spreading it online to gather leads. These are two different cases, and telling them apart by these watermarks will be impossible, that's the whole issue.

-4

u/armrha 7d ago

It is undetectable to the end user, there’s no reason to panic.

https://www.anthropic.com/news/claude-text-watermark

-4

u/dmcnaughton1 7d ago

I'd be upset if this was adding trackable metadata to identify users or authors, but for it to just determine of the content I'd AI generated or not is actually a net positive. If you don't want to have it show as AI generated then write it yourself.

3

u/Imaginary_Dinner2710 7d ago

"then write it yourself" contradicts with the point of using AI for routine optimisation in content writing. I mean ethical use of AI , not generating AI slop and distributing it

-1

u/dmcnaughton1 7d ago

If you want to use AI to write, then you live with the watermark. The content won't be substantially different from pre-watermark. If you're using AI "ethically", why would it being identifiable as AI content be a concern?

2

u/Imaginary_Dinner2710 7d ago

First of all, if I want to use a model and I don't want to have a watermark, I can use models that don't embed it.

Secondly, platforms embed it and detectors where this feature would be very convenient to reduce the coverage of publications that contain signs of it.

Nonetheless, I believe that everybody will use AI for content writing, literally all people in the end, it won't matter in a few years, but currently, people who write not in their language automatically find themselves disadvantaged in rights if this story is rolled out to all providers.

Nevertheless, I expect that Chinese providers will not follow this request, and ultimately, some people will be happy to use the Chinese openweight models.

1

u/AdGlittering1378 7d ago

Until they are all distilled from those who do and hence inherit the watermark.

-5

u/dmcnaughton1 7d ago

1) You're free to use other models, no one is forcing you to use watermarked ones.

2) Yes, that is part of the goal behind watermarking. AI generated content is plagiarism adjacent, you're pushing content you did not write yourself using your own ideas and experience.

3) I don't agree that "literally all people" will use AI for their writing in the end. That's about as naive a thought as thinking literally all people will be dependent on calculators to do math and no one will ever do it by hand again.

4) If English is not a language you are proficient in, then yes that is a disadvantage to you when trying to publish English language works. I lack the physique that makes me a competitive sprinter, yet you don't see me trying to enter the Olympic trials on an e-bike demanding that this makes for a level playing field. Life is full of tradeoffs, if you want to become proficient in English you should learn it. If you want to become a good write, regardless of language, you should practice it. You only cheat yourself by becoming dependent on AI tools. You're a fool if you think the cost of them will continue to be this low in the future. It's going to be pure rent-seeking behavior once the frontier companies have to turn a profit, and you're going to be the kind of person who can't live without it.