Source-led article
Claude text watermark: what it can and cannot prove about authorship

Anthropic has announced plans to mark text generated by its Claude models, a move designed to bring the company in line with Article 50 of the EU AI Act. The announcement, made this week, drew immediate reactions from writers who use Claude to edit their own work and objected to that work carrying an AI-generated label, and from developers who raised a separate set of concerns.
Forbes reported that people object to their own work becoming detectable as AI-assisted. TechCrunch found similar feelings on Reddit, alongside users arguing back at the complainers. Decrypt reported that open-source projects are emerging to bypass or disrupt the watermarks, and some reports suggest the marking system is already in use.
What Anthropic has confirmed
Anthropic’s help page states that models launched in the EU from August 2 onward will support marking at launch, and that work on earlier models is in progress. The page names no model that currently carries a mark. Forbes notes that users cannot opt out, and the help page does not mention that option either.
The company says marking will cover output from supported models worldwide, across its API, apps and developer tools. Anthropic will assist users and external parties in identifying its marks and plans to release technical details later.
What a detected watermark does and does not prove
A detected mark indicates that Claude’s systems processed the text at some point. It does not prove that Claude was the original author.
Users often proofread, translate, summarise or convert files through Claude. Any of these actions can introduce a mark even when the ideas and words originated with the human user. Content may also change after Claude processed it, which raises further questions about what the mark actually refers to.
The reverse problem matters just as much. Anthropic lists five reasons why marked content might not show a detectable mark, so the absence of a watermark is not evidence that text was written by a human.
The EU rules behind the move
Article 50(2) of the EU AI Act requires providers of generative AI systems to mark outputs such as audio, images, video and text in a machine-readable way to indicate that they are artificially generated or manipulated. Providers must make those marks effective, interoperable, robust and reliable, as far as that is technically possible.
Article 50(4) adds a separate obligation to label deepfakes and AI-generated text published as public-interest information. The European Commission has clarified that deployers cannot rely solely on the provider’s machine-readable mark to fulfil their own disclosure duties.
The marking duty does not apply where a system only assists with standard editing or does not substantially alter the input data or its meaning. The Commission’s final guidelines, published on July 20, list AI-generated translations alongside grammar correction and spellchecking as examples covered by this exception — yet Anthropic states that translated output can still bear a Claude mark. A detected mark, in other words, does not necessarily mean the mark was legally required.
The Code of Practice on Transparency of AI-generated Content offers a voluntary framework that about 190 organisations had signed by the end of July. Signatories in Section 1 include Google, Meta, Microsoft, OpenAI and Anthropic. Google, which says SynthID marks text generated through the Gemini app and web experience, signed on July 24 and has voiced concern that adding more rules while the technology is still evolving could undermine Europe’s competitiveness goals.
Short texts and weaker marks
The code applies watermarking to free-form text longer than 200 tokens and calls anything shorter “very short text”, expecting the cutoff to decline as methods improve. For audio, images, video and text in files circulated online, the code generally requires two separate marks because no single technique meets all four requirements. Free-form text cannot carry metadata, so the code accepts a single watermark layer for this format while acknowledging it may be less reliable than for longer passages.
What remains unclear
Anthropic has not described its implementation in technical detail, and neither of the two watermarking research papers presented at ICML in 2024 and 2025 tested it. That research still raises questions about durability. At ICML 2024, researchers from ETH Zurich’s SRI Lab showed that querying a watermarked model through its public API allows an attacker to infer enough about the scheme to remove or spoof marks, at a cost under $50 and an average success rate above 80%. Separate work found that at least 74% of good paraphrases of non-watermarked material were detected as watermarked, with an expected false-positive rate
Datos clave
| Punto | Detalle |
|---|---|
| Fuente | Search Engine Journal |
| Fecha | 2026-08-14T19:00:39+00:00 |
| Tema | What A Claude Watermark Can & Can’t Tell You About Authorship via @sejournal, @MattGSouthern |