top of page

Claude Is Watermarking Its Writing Here’s What That Actually Means

  • 15 hours ago
  • 4 min read

Anthropic has announced that it’s adding an invisible watermark to Claude’s writing. Unlike the watermark you see across a stock photo, however, you won’t be able to see the one in Claude’s text output. Instead, the watermark is built into the words Claude “chooses” when it writes. With the right detection tool, those choices can later provide evidence that Claude was involved in producing a particular piece of text.

Anthropic is introducing the system largely because of the EU AI Act, whose new transparency requirements came into effect on August 2. New Claude models will include watermarking from launch, while Anthropic plans to extend it to existing models over the coming months. Despite the requirement originating from European regulation, Anthropic says it will initially apply the system globally.

How Do You Watermark Words?

Large language models do not write sentences in the same way humans do. At each step, Claude is choosing between possible next words or, more precisely, tokens. Often there are several choices that would produce an equally sensible or “correct” sentence.

The watermarking technology takes advantage of those choices.

Anthropic describes a system in which the randomization involved in choosing between possible words is modified according to a secret key. To a reader, the result should look completely normal. The crucial part is that, across a sufficiently long passage, Claude's choices create a statistical pattern that someone with Anthropic's detection system can recognize.

In other words, Claude is not sneaking a secret message into your document. The statistical pattern in the writing itself is the watermark.

That also means copying and pasting the text does not remove it. Light editing may not either. A human completely rewriting the passage probably will, but at that point the question of whether the resulting text is still AI-generated becomes considerably murkier.

Anthropic is an early adopter, but it won’t be unique for long. The method used for Claude is a version of SynthID-Text based on research published by Google DeepMind in 2024. Because of the requirements of the EU AI Act, other companies will likely introduce their own watermarking systems soon.

A Watermark Is Not an AI Detector

Anthropic itself acknowledges an important limitation with the system. Despite what the term watermark might imply, it cannot reliably answer “Was this written by AI?” What it can answer is something closer to “How likely is it that Claude was involved in producing this text?

Detection also becomes harder with short passages and text where Claude has less freedom over word choice. Computer code is an obvious example: there are fewer alternative ways for Claude to express the same thing without changing how the code works. Highly factual writing presents a similar problem. Anthropic therefore expects the watermark to be weaker or more ambiguous in these situations.

An absence of a detectable watermark does not prove that something was human-written. As mentioned earlier, actions like heavy rewriting, mixing Claude's output with other text, or simply having too little text to analyze can make detection fail.

The Strange Problem of AI-Assisted Writing

AI-assisted writing is rarely as simple as “written by a human” or “written by AI”. Someone might write a paragraph and ask Claude to make it clearer. Someone else might give Claude a rough outline and then heavily edit what comes back. Another person might write something entirely in Swedish and ask Claude to translate it into English.

There is a wide spectrum between “I wrote this” and “an AI wrote this”, with much of today’s AI-assisted work sitting somewhere in the middle.

The problem is that a watermark cannot tell you where on that spectrum a particular piece of writing belongs. It can provide evidence that Claude was involved in producing the text, and generally the more opportunity Claude has to choose the words, the stronger the statistical signal can become. What the signal cannot reveal is what that collaboration actually looked like.

This distinction matters because simply detecting Claude’s involvement can easily be interpreted as something much stronger: that Claude wrote the content. The colleague who uses Claude to improve an argument and the colleague who asks Claude to generate one from scratch have used the technology in fundamentally different ways, even though both pieces of writing could contain evidence of Claude’s involvement.

“AI was involved” and “AI did the work” are not the same statement. If we start treating them as interchangeable, a technical signal designed to provide transparency could quickly become a judgment about whether someone’s work is genuinely their own.

What Happens Next?

Anthropic plans to make detection available to users and third parties, although the detection system is not yet generally available. The company is also adding digitally signed records to supported files that can verify their origins and whether AI was involved, using standards such as C2PA (Coalition for Content Provenance and Authenticity).

The key result here is that watermarking will make it much harder for organizations to use AI without explicitly disclosing its role in the finished work. Companies, publishers and institutions may increasingly have to come clean about where and how AI enters their work — not because the technology is necessarily doing anything wrong, but because its involvement is becoming detectable.

That forces a harder conversation about what we might call AI shaming. If using Claude to sharpen an argument or rewrite a paragraph leaves a detectable trace, does that make the work somehow less legitimate? We are rapidly normalizing AI as a tool while still treating its use as something people are expected to confess. Watermarking may make AI use more transparent, but what it really ought to do is force us to decide whether using AI is something worth being ashamed of in the first place. 

AI watermarking is a technical solution to a very human problem related to culture and ways of working. Organizations today need to decide what responsible AI use should actually look like. Foundational questions like “when is it okay for us to use AI?”, “when should AI use be disclosed?”, and “what norms should we follow around AI assistance?” still need answering regardless of how foolproof watermarking becomes. Questions like these are the ones we work on every day at Stellar Capacity to help organizations figure out what AI adoption actually means in practice.

This article was researched and written with the assistance of Generative AI tools.

Contact us if you would like to know more about our programs and one of our program advisors will get in touch!

Thank you!
bottom of page