Field Notes
The Watermark Hid Between Two Good Words
A statistical watermark is strongest where a model had room to choose. That makes it evidence of discretion, not authorship, truth, or responsibility.
A sentence reaches a small fork.
The afternoon was cold and gray. It might just as easily have been cold and overcast. Both words fit. Neither changes the weather. The choice feels like the least consequential part of the paragraph.
That is where a watermark can live.
Anthropic recently explained how future Claude models will mark generated text. There are no hidden characters and no extra line of metadata tucked beneath the prose. The system changes the source of randomness used when the model has several acceptable next words. A private key and the preceding text help choose among those words. Repeated across a long enough passage, the choices form a statistical pattern that a detector with the key can recognize.
The mark is not attached to the sentence after writing. It happens while the sentence is deciding how to sound.
This is a clever response to a real policy demand. The European Union's new transparency code asks providers to make generated text machine-readable and detectable where technically feasible. Anthropic is using a version of Google DeepMind's SynthID-Text method, whose published evaluation found no measurable quality loss in benchmarks, side-by-side ratings, or a live test involving nearly twenty million Gemini responses.
But the technical details give the watermark an interesting shape.
It is strongest where the model had room to move.
A long, open-ended explanation offers many low-stakes choices. A factual sentence offers fewer. Once the model writes that two plus two equals something, only one ordinary answer belongs next. Code has names, syntax, and behavior that often make one token materially better than another. A light proofreading pass may change too few words to leave a detectable pattern. A translation, by contrast, can carry a strong mark because the model chooses nearly every word in the new language.
The watermark records discretion.
That is not the same as recording authorship.
A person can spend days forming an argument, gathering interviews, and drafting a report, then ask Claude to translate it. The resulting text may carry a strong statistical signal even though the reporting, structure, and ideas came from the person. Another person can ask for a complete code patch and receive an output with a weaker mark because the code allowed fewer harmless substitutions. The detector may see more machine involvement in the translation than in the delegated implementation.
It would be technically correct and socially easy to misunderstand.
Anthropic names this limit directly: a positive result can indicate that Claude was probably involved, but cannot distinguish between writing and heavy editing. It says nothing about ownership or legal responsibility. A complete rewrite can remove the signal. Short samples may not contain enough choices to judge at all.
Those limits need to become part of the detector interface, not footnotes people discover after a dispute.
Imagine the result appearing in a classroom, newsroom, hiring process, or workplace review. A bright badge reading “AI-generated” would convert a probabilistic trace into a verdict about how the work was made. It would invite a manager or teacher to infer intention, effort, and honesty from a mechanism that knows none of those things. The watermark does not know whether the model supplied the thesis or the commas. It only knows that enough word choices resemble the keyed pattern.
A responsible result would be less satisfying. It would say which provider's mark was checked, how much text was available, how confident the detector is, and when it has abstained. It would describe the editing and translation limits before anyone uses the result as evidence against a person. It would offer an appeal path that asks for drafts, notes, version history, or process disclosure instead of treating one score as the whole history of the work.
This sharpens the argument in Provenance Became Part Of The Interface. Provenance is a record of the path an artifact took: who initiated it, what tools acted, what changed, and who accepted responsibility. A watermark is narrower. It can be one useful witness to model contact, but it did not watch the work happen.
The difference is even clearer beside a content credential. The C2PA specification can bind signed claims about a media file's origin and edits to the asset. It explicitly avoids judging whether those claims make the content good or bad. A text watermark carries less history but can survive where file metadata would be stripped away. Neither one can decide whether the argument is true, the idea is original, or the person using the tool acted responsibly.
We should still want marks that make synthetic media easier to study and large-scale deception harder to hide. A provider-held signal is more grounded than a detector guessing from tidy paragraphs and suspiciously enthusiastic em dashes. The mistake would be asking detectability to become a moral category.
The watermark hid between two good words because either word would do.
Responsibility begins where either word is no longer the whole story.