On October 5, 2026, OpenAI announced that ChatGPT and Codex will start embedding invisible watermarks into their text output for EU users. The feature is called textGrain, and it exists because of one law: the EU AI Act.
If you build on top of OpenAI's models, or you're just curious how "invisible AI watermarking" actually works under the hood, here's a breakdown.
Why now
The EU AI Act's transparency rules took effect on August 2, 2026. They require AI companies to mark AI-generated content in ways that can be detected, not necessarily by the naked eye, but by a tool built for the job.
OpenAI put it plainly: "The EU AI Act's transparency rules require AI companies to mark AI-generated content."
This isn't OpenAI being proactive out of goodwill. It's compliance. Which, if you've followed this space, is worth noting, because OpenAI has had a working text watermarking technique sitting on the shelf for roughly a year, and had previously chosen not to ship it over concerns about bias against non-native English speakers and how easily it could be defeated. The EU AI Act appears to have forced the decision.
Anthropic shipped something similar for Claude back in August, for the same regulatory reason. This is quickly becoming table stakes for any model provider operating in the EU.
How textGrain actually works
Here's the part developers usually want to know: it doesn't touch pixels, metadata, or hidden Unicode characters. It lives in the words themselves.
From OpenAI's own explanation, the watermark works by subtly shaping the model's word choices during generation. Every time the model predicts the next token, there's usually a cluster of statistically similar options it could pick. textGrain uses a secret key to quietly bias which of those near-equal choices gets selected, across hundreds of these micro-decisions in a single response.
The result: a pattern invisible to a human reader, but detectable by a tool that knows the key.
A rough mental model, if you've worked with token sampling before:
for each token position:
candidates = model.get_top_k_predictions()
# normally: pick highest probability candidate
# with textGrain: use secret_key to bias selection
# among near-equal-probability candidates
chosen_token = watermark_bias(candidates, secret_key)
No visible artifact, no change in meaning or quality, just a statistical fingerprint spread across the whole response. Multiply that pattern across enough words, and a detector can pick it out with reasonable confidence.
Importantly, OpenAI says the method doesn't identify individual users. It's a content-level signal, not a tracking mechanism tied to your account.
Where it breaks down
This is the part that matters if you're evaluating how reliable this actually is:
- Editing kills it fast. Replacing just 10% of the words in a watermarked response dropped detection accuracy from 92% down to 66%. A light paraphrase pass, and the signal gets noisy quickly.
- Short text is hard to watermark reliably. There aren't enough token decisions in a two-sentence answer to build a strong statistical pattern.
- Math answers and translated text are harder to detect. Both involve more constrained or transformed token choices, which weakens the fingerprint.
- A missing watermark proves nothing. If a detector doesn't find the pattern, that does not mean the text wasn't AI-generated. It could just mean the text was edited, translated, or too short.
In short: textGrain is a detectable-pattern tool for a specific regulatory requirement, not a forensic-grade proof system. Treat any "this text is/isn't AI-generated" claim built on it with appropriate skepticism.
Scope and rollout
- Products: ChatGPT and Codex
- Region: Rolling out to EU users across all plans over the coming weeks
- API: Developers worldwide can enable it via the API, but it's off by default
- Detector access: Limited to approved researchers for now, not a public tool
That last point matters if you were hoping to verify AI-generated text yourself. You can't, at least not yet. OpenAI is keeping the detection side gated, likely both to prevent adversarial reverse-engineering of the watermark and to control how the "is this AI-written" question gets answered publicly.
What this means if you're building on GPT
If your product uses the API to generate text that EU users will read, this is worth understanding even though it's opt-in today. Two things to keep on your radar:
- Regulatory direction is clear. Watermarking AI output is becoming a compliance baseline across major providers (OpenAI, Anthropic, likely Google and Meta following the same Code of Practice). If you're building for EU markets, budget time to understand how this affects your product, even if you're not required to enable it directly.
- Don't build detection logic around it yet. With detector access restricted and accuracy dropping sharply after light editing, textGrain isn't something you can reliably use today to verify content provenance in your own pipeline.
This is a compliance feature first, and a content-provenance tool a distant second. Worth watching how it evolves as more providers follow the same path.
Top comments (0)