PaPoo
cover

Watermarks Won’t Save Us From the Real Claude Problem

What surprised me here is not that Anthropic is adding watermarking. It’s that anyone is still talking about this as if detection were the hard part. The hard part is governance, and watermarking only scratches that surface. If a model can leave a trace through copy-paste and some light editing, that’s useful. If it can be stripped by paraphrasing, translating, or folding Claude output into a larger draft, then you already know where the edge cases are. That’s most of the real world.

I do think this is a sensible move for Anthropic, especially if the EU AI Act is pushing them in that direction. For developers building with Claude, the practical impact is probably less dramatic than the headline makes it sound. If you’re generating product copy, internal docs, support replies, or draft prose, a watermark that survives simple reuse could help downstream teams decide when to ask questions. That’s not nothing. But it also won’t give anyone the magical “this was definitely written by Claude” button that institutions keep wishing for.

The article’s most interesting line is the least glamorous one: finding a watermark doesn’t prove Claude authored the text. That matters a lot. A lot of people will hear “watermark” and immediately imagine a clean forensic signal. That’s not what this is. It sounds more like a weak provenance hint, useful in some cases and misleading in others. If a student uses Claude to proofread or translate something, the mark may still be there. So the detection story can cut both ways. That’s where I’d be cautious about schools and publishers treating this as dispositive.

This also feels like a reminder that the “AI novel written undetected” era was always going to be fragile. Not because of some grand technical breakthrough, but because the ecosystem is slowly making the easy path more legible. Still, I wouldn’t overread the competitive angle. Google DeepMind already has its own watermarking story with SynthID. Anthropic isn’t inventing the category; it’s catching up to a pressure that was already there.

What I’d actually want to see next is not more branding around watermarking, but better tooling around chain-of-custody. Who generated the text, with what model, in what workflow, and what got edited afterward? That’s the part publishers, universities, and teams building on Claude really need. A watermark is one clue. It is not a verdict.


Reference: Anthropic rolled out a feature that stops undetected AI-generated writing

同じ著者の記事