← back to the archiveCover illustration for “Claude's watermark is provenance, not a verdict”
ESSAYday 73·4w ago·by Andy Padia

Claude's watermark is provenance, not a verdict

Claude's watermark can show that Claude processed text, not who authored it or whether policy was breached. Route detections to review, never straight to discipline.

An employee document trips Claude's text watermark. That is enough to ask a question. It is nowhere near enough to decide that the employee outsourced the work, breached policy or even asked Claude to write it.

On August 11, Anthropic documented its machine-readable marking plan for models launched in the EU from August 2. The Claude marking page, in its August 12 version, said text from supported models would carry an imperceptible watermark, supported files would carry signed C2PA provenance metadata, older models were still in transition, and general detector documentation had not yet shipped. More important than the launch state was the limitation Anthropic put beside it: finding a mark and failing to find one are both inconclusive.

The mark records contact, not authorship

A detected Claude mark can mean Claude generated the text. It can also mean Claude proofread, translated, summarised or converted material whose ideas and words came from somewhere else. The content may then have been edited, excerpted or combined again.

The opposite inference breaks too. Anthropic lists heavily edited or translated text, mixed authorship, very short passages, models released before marking was supported, and unsupported products or file types as reasons a Claude-processed output might not carry a detectable mark. For files, re-saving, screenshots or format conversion can strip metadata.

So the detector does not produce “AI” and “human.” Its defensible outputs are smaller: supported mark found and supported mark not found. My operating rule is that either result may route a review; neither result can close one.

This is the same category error I wrote about in Substack's optional AI detector. That product risked turning an uncertain technical estimate into a social judgment about cooperation. Claude's watermark is a stronger provenance mechanism, but strength does not change what the mechanism measures.

The law does not turn it into an employee test

Article 50(2) of the AI Act puts a machine-readable marking and detection duty on providers of systems that generate synthetic audio, image, video or text. It also excludes standard assistive editing, or processing that does not substantially alter the deployer's input or its meaning, from that duty.

The European Commission's June 10 explanation says the transparency rules cover provider marks, chatbot notices, deepfakes and certain public-interest text labels. It does not define a watermark as proof of employee authorship, misconduct, ownership or work quality.

An enterprise may still have a legitimate reason to verify marks. The mistake is importing a provider's transparency control into an HR decision without changing the evidentiary standard. Compliance with a marking duty and proof of a workplace breach are different jobs.

Put the policy through the two-document test

Here is an illustrative policy test, not a client incident. An engineer writes a design note and uses Claude only to translate it; the mark survives. Another submits a largely generated note from an older model, then edits it until the signal disappears. A rule that says “detected means breach” punishes the permitted workflow. A rule that says “not detected means clear” misses the prohibited one.

That is not the detector failing. It is the organisation asking provenance to make a policy judgment.

I would not allow a watermark result to become the sole ground for discipline, performance assessment or an authorship claim. Pair it with the employee's declared use, the policy that applied at the time, the model and surface used, revision history, lawful audit records where they exist, and a human review with an appeal path.

A detected Claude mark and no detected mark both route to workflow evidence before a human-owned decision.

The diagram is deliberately asymmetric. A detected mark supports one narrow statement: Claude may have processed the content. No mark leaves origin unknown. The useful decision evidence sits downstream of both branches.

Detection access is a separate governance problem

Who gets the detector matters as much as what it returns. An August 4 watermarking preprint defines three risks from broadly delegated detection keys: watermark sanitisation, scope abuse and user profiling. Its authors propose attribute-bound keys that can only test content matching an authorised policy.

That paper describes a separate research prototype, not Anthropic's undisclosed implementation on August 12. I am using its threat model, not claiming its mechanism sits inside Claude. The procurement questions still travel: who can scan, for which purpose, at what volume, with what logging and retention, and whether the person affected can challenge the inference.

A detector issued for regulatory verification should not quietly become a batch scanner for employee documents. Without purpose limits, access logs and review, a transparency feature becomes workforce surveillance with a cryptographic-looking answer.

Store the uncertainty, not a fake binary

The record I would put into a content or policy workflow has three states: detected, not detected and not testable. Beside that, store the model family and version when known, the date, the product surface, any translations or edits, the declared use, and the human-owned disposition. If the detector or policy changes, keep the version used for the decision.

On September 10, the evolving Anthropic page had been updated to describe private-preview detector access and named supported models. That is later state, not evidence available on August 12. I did not test the detector, and Anthropic had not published false-positive rates, false-negative rates, a minimum reliable passage length or editing-robustness measurements in the August 12 documentation.

The missing numbers do not make the watermark useless. They define its proper place: one provenance input inside a reviewable process, never the verdict at the end of it.

A Claude watermark can tell you the tool touched the work; only the workflow evidence can tell you what that contact meant.

#claude#watermarking#provenance#ai-governance#eu-ai-act#employee-policy
← older drop
QM makes the coding agent a swappable part
newer drop →
Agent autonomy is a timeout setting, not a capability

related drops

explore all 128 drops →
← back to the archiveday 105