Most of us think of a watermark as something you can see.
A logo in the corner. A faint stamp across a photo.
Text doesn't have a corner.
So how do you mark a paragraph without changing how it reads?
That question stopped being academic this week.
- On October 5, 2026, OpenAI published Our approach to EU text provenance rules. API customers anywhere can now opt in to text watermarking for select models. It stays off by default. Over the coming weeks, eligible ChatGPT and Codex output in the European Union gets an invisible watermark. The detector is not public. Approved researchers and expert organizations can apply for access.
- The same day, OpenAI posted the textGrain technical report. It says the detector "requires only the generated text and the secret key."
- OpenAI's help center adds that textGrain "does not add hidden characters, invisible spaces, or unusual punctuation." The signal lives in which words get picked.
- The reason is regulation. The EU AI Act's transparency rules took effect on August 2, TechCrunch reported, and they require AI-generated content to be marked in a way machines can identify.
Here's the part I find most interesting. OpenAI published its own failure modes next to the launch.
At a 1% false positive target, its detector found the watermark in about 95% of 400-token passages and about 80% of 200-token passages, for content like psychology. Math was "substantially lower." And in a 400-token test, replacing 10% of the words with synonyms dropped detection from about 92% to 66%. Replacing 25% dropped it to 17%.
Then this line:
"The absence of a detected watermark does not prove human authorship."
That distinction matters. A watermark is evidence when you find it. It is not evidence when you don't.
I wanted to feel that in code. So let's build a tiny version.
By the end, you'll run one command:
npx tsx watermark.ts
And watch a keyed watermark go into ordinary looking sentences, get detected with only the text and the key, and then fade as we swap synonyms.
No API key.
No model.
Just TypeScript.
One honesty note: this is not textGrain. It's the classic textbook version of the same family: keyed random numbers steer sampling, and the detector adds up scores and compares them to a Gamma(n, 1) distribution, the same test shape the textGrain report describes. There's no optimal transport, no entropy budget and no real tokenizer. The "model" is a mock with a seven-slot vocabulary. Every number in the output comes from my toy, not from OpenAI.
Table of Contents
- What We Are Building
- Project Setup
- Step 1: A Tiny Mock Model
- Step 2: Keyed Randomness
- Step 3: Sampling, With and Without a Key
- Step 4: The Detector
- Step 5: The Edit Attack
- Step 6: Run It
- Where It Breaks Down
- The Bigger Idea
What We Are Building
A language model writes one token at a time. At each step it has a probability for every option and rolls a die.
A watermark replaces the die.
Instead of fresh randomness, the model uses a random number computed from a secret key and the last few words. Over many keys, the choices land with exactly the model's probabilities. So the text reads the same.
But someone holding the key can recompute those numbers. Watermarked text keeps picking words whose keyed numbers were unusually high. Normal text doesn't.
secret key + last 3 words + candidate word
↓
keyed number r in (0, 1)
↓
pick the word with the largest r ^ (1 / p)
↓
later: detector recomputes r for each word it sees
↓
score too high to be luck? watermark
Three pieces:
- A generator that samples with the key.
- A detector that needs the key and the text. Nothing else.
- An attacker who swaps words for synonyms.
Project Setup
mkdir tiny-text-watermark && cd tiny-text-watermark
npm init -y
npm install -D tsx typescript @types/node
Save the following TypeScript blocks in order as watermark.ts. It only uses node:crypto.
Step 1: A Tiny Mock Model
Our "language" is a sentence pattern with seven slots. Each slot holds words that all fit.
import { createHmac } from "node:crypto";
// A tiny "language": each slot is a group of words that all fit there.
const SLOTS: string[][] = [
["The", "Our", "This", "That"],
["small", "tiny", "simple", "quick", "careful", "new"],
["agent", "model", "service", "worker", "tool", "system"],
["fixed", "patched", "repaired", "checked", "reviewed", "updated"],
["the", "one", "each", "every"],
["bug", "issue", "test", "query", "report", "file"],
["quickly.", "calmly.", "carefully.", "safely.", "quietly.", "today."],
];
type Dist = { word: string; p: number }[];
// A deterministic 32-bit hash, so the toy model is the same every run.
function hash(s: string): number {
let h = 2166136261;
for (let i = 0; i < s.length; i++) {
h ^= s.charCodeAt(i);
h = Math.imul(h, 16777619);
}
return h >>> 0;
}
// The mock model: a next-word distribution for slot t, given the last word.
// "open" text has many good choices. "rigid" text (think math) has one.
function nextWordDist(t: number, prev: string, style: "open" | "rigid"): Dist {
const group = SLOTS[t % SLOTS.length];
if (style === "rigid") {
const top = hash(prev) % group.length;
const rest = 0.1 / (group.length - 1);
return group.map((word, i) => ({ word, p: i === top ? 0.9 : rest }));
}
const weights = group.map((w) => 1 + (hash(prev + "|" + w) % 8));
const total = weights.reduce((a, b) => a + b, 0);
return group.map((word, i) => ({ word, p: weights[i] / total }));
}
The important part is nextWordDist(). It plays the model.
In "open" style, the probabilities shift with the previous word, and there are several good choices at every slot. That's prose.
In "rigid" style, one word gets 90%. That's math, or code, or any text where the next token is mostly forced.
Hold onto that. It comes back.
Step 2: Keyed Randomness
// Keyed randomness: a number in (0, 1) for (context, candidate word).
// Only someone with the key can recompute it.
function prf(key: string, context: string, word: string): number {
const mac = createHmac("sha256", key).update(`${context}\u0000${word}`).digest();
return (mac.readUInt32BE(0) + 0.5) / 2 ** 32;
}
// The context is the three previous words.
function contextAt(words: string[], t: number): string {
return [t - 3, t - 2, t - 1].map((i) => words[i] ?? "^").join(" ");
}
prf() turns the key, the context and a candidate word into a number between 0 and 1 with HMAC-SHA256.
Without the key, these numbers look like noise. With the key, anyone can recompute them later.
The context is the last three words. That's what lets the detector work on text alone, with no prompt and no model.
Step 3: Sampling, With and Without a Key
Normal sampling first.
// A seeded random number generator, so every run prints the same numbers.
function rng(seed: number): () => number {
let a = seed >>> 0;
return () => {
a = (a + 0x6d2b79f5) >>> 0;
let x = Math.imul(a ^ (a >>> 15), 1 | a);
x = (x + Math.imul(x ^ (x >>> 7), 61 | x)) ^ x;
return ((x ^ (x >>> 14)) >>> 0) / 2 ** 32;
};
}
// Normal sampling: roll a die weighted by the model's probabilities.
function sample(dist: Dist, rand: () => number): string {
let u = rand();
for (const { word, p } of dist) {
if ((u -= p) <= 0) return word;
}
return dist[dist.length - 1].word;
}
Now the watermarked version.
// Watermarked sampling: replace the die with keyed randomness.
// Pick the word with the largest r ** (1 / p). Averaged over keys,
// this picks each word with exactly probability p.
function sampleWatermarked(dist: Dist, key: string, context: string): string {
let best = dist[0].word;
let bestScore = -Infinity;
for (const { word, p } of dist) {
const score = Math.log(prf(key, context, word)) / p;
if (score > bestScore) {
bestScore = score;
best = word;
}
}
return best;
}
type Options = { tokens: number; style: "open" | "rigid"; key?: string; seed: number };
function generate({ tokens, style, key, seed }: Options): string[] {
const rand = rng(seed);
const words: string[] = [];
for (let t = 0; t < tokens; t++) {
const dist = nextWordDist(t, words[t - 1] ?? "^", style);
words.push(key ? sampleWatermarked(dist, key, contextAt(words, t)) : sample(dist, rand));
}
return words;
}
The most important line is Math.log(prf(...)) / p.
Picking the largest r ^ (1 / p) is an old trick. If r were truly random, it would choose each word with exactly probability p. So the watermark doesn't bend the model's preferences. It just makes the choice reproducible for whoever holds the key.
That's why OpenAI can say it saw no meaningful change in benchmark scores with the watermark on. The distribution is the same. Only the source of randomness changed.
Step 4: The Detector
// P(Gamma(n, 1) >= s), the chance of a score this high with no watermark.
function gammaTail(n: number, s: number): number {
if (n === 0) return 1;
let logTerm = -s; // log of e^-s * s^k / k!, starting at k = 0
let logSum = logTerm;
for (let k = 1; k < n; k++) {
logTerm += Math.log(s) - Math.log(k);
const hi = Math.max(logSum, logTerm);
logSum = hi + Math.log(Math.exp(logSum - hi) + Math.exp(logTerm - hi));
}
return Math.min(1, Math.exp(logSum));
}
type Verdict = { scored: number; score: number; pValue: number; detected: boolean };
// The detector needs only the text and the key. No model, no prompt.
function detect(words: string[], key: string, alpha = 0.01): Verdict {
const seen = new Set<string>();
let score = 0;
for (let t = 0; t < words.length; t++) {
const context = contextAt(words, t);
if (seen.has(context)) continue; // score each context once
seen.add(context);
score += -Math.log(1 - prf(key, context, words[t]));
}
const pValue = gammaTail(seen.size, score);
return { scored: seen.size, score, pValue, detected: pValue < alpha };
}
For every word, the detector recomputes r for the word that was actually chosen and adds -log(1 - r).
With no watermark, each term averages 1. So the total for n words follows a Gamma(n, 1) distribution.
With a watermark, chosen words have suspiciously high r. The total runs hot.
gammaTail() turns the total into a p-value: the chance of a score this high from text that was never watermarked. Below 1%, we call it.
Two details matter:
- Score each context once. If the same three words repeat, the same keyed numbers repeat. Counting them twice would fake evidence. The textGrain report has the same rule.
- The detector never sees the model. No prompt, no probabilities. Just text and key.
Step 5: The Edit Attack
// The edit attack: swap a share of words for a synonym from the same slot.
function swapSynonyms(words: string[], share: number, seed: number): string[] {
const rand = rng(seed);
return words.map((word, t) => {
if (rand() >= share) return word;
const others = SLOTS[t % SLOTS.length].filter((w) => w !== word);
return others[Math.floor(rand() * others.length)];
});
}
// What a detection result can and cannot say.
function report(v: Verdict): string {
if (v.detected) return `watermark detected (p = ${v.pValue.toExponential(1)})`;
return `no watermark detected (p = ${v.pValue.toFixed(2)}). This does not prove a human wrote it.`;
}
swapSynonyms() is the cheapest attack there is. Pick some words, replace each with another word from the same slot.
Each swap breaks two things: the swapped word's own score, and the context for the next three words.
report() is the part I'd ship first. A positive result says what it found. A negative result says what it can't prove.
Step 6: Run It
const KEY = "demo-key-not-a-secret";
const show = (w: string[]) => w.join(" ").replace(/\. (\w)/g, (_, c) => `. ${c}`);
console.log("One passage, 21 words\n");
const marked = generate({ tokens: 21, style: "open", key: KEY, seed: 1 });
const plain = generate({ tokens: 21, style: "open", seed: 1 });
console.log("watermarked: ", show(marked));
console.log(" ->", report(detect(marked, KEY)));
console.log("unwatermarked:", show(plain));
console.log(" ->", report(detect(plain, KEY)));
console.log("wrong key: ", report(detect(marked, "some-other-key")));
const TRIALS = 500;
// Each trial gets its own key, so every passage is different.
type Make = (tokens: number, key: string, seed: number) => string[];
function detectionRate(tokens: number, make: Make): string {
let hits = 0;
for (let i = 1; i <= TRIALS; i++) {
const seed = tokens * 10_000 + i;
const key = `${KEY}-${seed}`;
if (detect(make(tokens, key, seed), key).detected) hits++;
}
return `${((100 * hits) / TRIALS).toFixed(1)}%`;
}
const open = (n: number, key: string | undefined, s: number) => generate({ tokens: n, style: "open", key, seed: s });
const rows: [string, Make][] = [
["no watermark (false positives)", (n, _k, s) => open(n, undefined, s)],
["watermarked, untouched", (n, k, s) => open(n, k, s)],
["watermarked, 10% of words swapped", (n, k, s) => swapSynonyms(open(n, k, s), 0.1, s)],
["watermarked, 25% of words swapped", (n, k, s) => swapSynonyms(open(n, k, s), 0.25, s)],
["watermarked, rigid text (math-like)", (n, k, s) => generate({ tokens: n, style: "rigid", key: k, seed: s })],
];
const lengths = [14, 28, 56];
console.log(`\nDetection rate over ${TRIALS} passages, 1% false positive target\n`);
console.log("".padEnd(36) + lengths.map((n) => `${n} words`.padStart(10)).join(""));
for (const [label, make] of rows) {
console.log(label.padEnd(36) + lengths.map((n) => detectionRate(n, make).padStart(10)).join(""));
}
Run it:
npx tsx watermark.ts
Real output from my machine:
One passage, 21 words
watermarked: Our quick service checked the test today. Our small model checked the query safely. This careful service checked the test today.
-> watermark detected (p = 9.1e-6)
unwatermarked: This small worker updated every test quietly. That simple system repaired each bug safely. The small service fixed one file quickly.
-> no watermark detected (p = 0.99). This does not prove a human wrote it.
wrong key: no watermark detected (p = 0.29). This does not prove a human wrote it.
Detection rate over 500 passages, 1% false positive target
14 words 28 words 56 words
no watermark (false positives) 1.0% 0.6% 1.0%
watermarked, untouched 95.6% 98.6% 99.4%
watermarked, 10% of words swapped 62.4% 81.0% 88.6%
watermarked, 25% of words swapped 25.8% 32.6% 43.2%
watermarked, rigid text (math-like) 18.4% 21.8% 18.2%
Read that table slowly. It tells the same story OpenAI's numbers tell, in miniature.
Clean watermarked text gets caught. 95.6% at 14 words, 99.4% at 56. Longer text gives the detector more evidence.
The false positive rate stays near the 1% target. That's the Gamma(n, 1) math doing its job.
Light edits hurt. Swapping 10% of the words takes the short passages from 95.6% to 62.4%.
Heavy edits mostly win. At 25%, most short passages slip through, and even 56 words only gets caught 43.2% of the time.
Rigid text barely carries a signal. When one word gets 90%, the keyed numbers rarely change the choice, so there's little to measure. And length doesn't rescue it here, because rigid text keeps repeating the same three-word contexts, and the detector only scores each one once.
Those are toy numbers from a toy model. The shape is the point.
Where It Breaks Down
Paraphrasing and translation. Synonym swaps are the gentle version. Rewrite the passage or translate it and the signal is gone. OpenAI's help center says the same for "substantial paraphrasing, or translation."
Short and rigid text. Short answers, math and code give the detector little to work with. OpenAI's help center notes the EU Code of Practice doesn't require watermarks in outputs under 200 tokens or in code snippets.
The key is the whole system. Whoever holds the key can detect, and with enough samples could probably forge or scrub. I think that's part of why detectors get gated, alongside the missed watermarks and false positives OpenAI names.
Many checks, many false alarms. A 1% false positive rate means roughly 1 in 100 unwatermarked passages gets flagged. Run a detector over a whole class or a whole inbox and someone innocent gets flagged.
Mixed authorship. A human draft with one AI paragraph, or AI text a human heavily edited. A watermark can say an OpenAI system touched part of a passage. In OpenAI's words, not "how much human judgment, editing, or creativity went into it."
No watermark is not a verdict. Other providers, older models, unsupported models, or the API with the setting off all produce clean text. Absence tells you almost nothing.
The Bigger Idea
Generation
↓ keyed choices (watermark)
Text leaves your system
↓ copy, paste, edit, translate
Detector
↓
"found a signal" or "can't say"
The model provides the words.
The key provides the signal.
The detector provides a probability.
None of them provides proof.
That's the shift I keep thinking about. Provenance is moving into the content itself. Images got Content Credentials and SynthID watermarks. Now text gets textGrain. Different formats. Same direction.
But the signal is weakest exactly where people want certainty: short answers, edited essays, code.
If you build with these models, don't wait for a public detector to tell you where your text came from. Log it where it's generated. A record in your own system beats a statistic recovered later.
A watermark is evidence when you find it. Its absence is just silence.
Software Should Explain Itself
Codex text output is part of OpenAI's EU rollout, which tells you where this is heading. More of your codebase is written with AI, and "where did this come from" is becoming a real question.
I'm building Helix to answer it from the engineering graph: who changed what, why it exists, and what it touches. Not a statistic. A record.
What critical engineering knowledge is your team losing right now?


Top comments (1)
Good.
I wanna have meaningful conversation about collaboration with you.
I believe we can achieve something big by collaboration.
How about you?