OpenAI's October 5 text-provenance announcement is a useful prompt for builders. Its textGrain watermark puts a statistical signal into eligible model output. API customers can opt in for select models; eligible ChatGPT and Codex text in the EU is due to receive it over the coming weeks. Detector access starts with approved researchers and expert organizations, not a public endpoint that every product can call.
The engineering lesson is broader than this rollout: an origin signal is one observation in a decision workflow. If your data model turns that observation into written_by_ai: true or written_by_human: true, it has already lost the uncertainty the source warns about.
Separate capture, signal, and decision
I would model a provenance review as three connected records, each with a different job.
Capture records what the workflow knew when the document was created: a version hash, the policy for permitted AI use, an author declaration, and references used for material claims. A hash identifies the version under review; it does not prove who wrote it. Keep a link to the version history where one exists.
Signal records what a check actually observed, including whether that check was applicable. A watermark observation needs its provider, method version, date, input version, and known limits. not_detected is a result about a check, not a statement of human authorship. not_run, unsupported, and inconclusive are useful states too.
Decision records what a person did with the evidence: which claims were checked, who reviewed them, whether the work was approved or needs revision, and why. The decision can remain pending when the evidence is incomplete.
Here is a proposed internal shape, not an OpenAI API response:
{
"document_version": "sha256:...",
"declared_ai_use": "unknown",
"watermark_check": {
"eligibility": "unknown",
"result": "not_run",
"method_version": null
},
"review": {
"source_check": "pending",
"decision": "pending",
"owner": null
}
}
This example records a check that has not run and a decision that has not been made. Other possible check results include detected, not_detected, and inconclusive. The exact fields will vary. The key property is that missing evidence stays missing. A downstream dashboard should never relabel not_detected as “human written,” or detected as misconduct.
Make the limits visible in the interface
OpenAI says a text watermark cannot identify the user, measure human contribution, establish ownership or responsibility, or verify accuracy. It also says that no detected watermark does not prove human authorship. Editing, translation, short passages, older output, and unsupported models can all complicate detection.
Its own evaluation shows why the qualification matters. For 400-token English passages in one test, replacing 10% of words with synonyms reduced detection from about 92% to 66%; replacing 25% reduced it to 17%. Those are company-reported results under specific conditions, not a field benchmark you can apply to every document.
Show that scope beside a result. An interface that displays only a red or green badge invites an overconfident decision. A review screen should show the input version, whether the method was eligible, the observed result, and what the result cannot establish. If a tool returns a probability or score, preserve its meaning and calibration rather than converting it into a verdict through a hidden threshold.
Route consequential cases to people
Consider a hypothetical policy memo. The team permits disclosed AI assistance and needs accurate claims plus a named approver. A detected watermark might justify asking how the memo was drafted. It cannot tell the team whether the memo's sources support its recommendations.
The review queue should bring the memo version, declaration, source links, and prior edits together. Reviewers can then check the claims that matter, ask for missing context, and record a decision. If the result could affect a person's reputation or opportunity, include a way to explain the evidence and correct errors. Do not make an adverse action the automatic output of a detector callback.
You can build this queue without a watermark detector. In fact, the current access limit means many teams must. Capture and review are still useful when a signal is unavailable, and they continue to matter if access expands.
Test the workflow, not just the classifier
If you can evaluate a provenance signal for your own authorized material, include clean, edited, short, translated, and constrained examples. Keep examples outside the tuning set. Measure false alarms, missed signals, inconclusive cases, reviewer time, and the decisions that followed. Separate results by content type; one overall rate can hide the cases your product handles poorly.
Write the proposed action for each result before looking at the scores. “Ask for supporting records” is a different consequence from “reject the work.” The evidence needed for those actions should differ too. Make it possible to pause when the method's scope is unknown.
My first implementation step would be a document-version record and a small review form with an explicit unknown state. Then I would test whether the team can answer who approved a claim and what source supported it. A detector can be assessed later as one additional observation. The system is useful as soon as it can preserve evidence and make uncertainty visible.
Top comments (0)