OpenAI starts watermarking ChatGPT's text in the EU -- and its own numbers show how easy it is to strip
"textGrain" steers word choices into a detectable pattern to satisfy the EU AI Act's labeling rule, following Anthropic's near-identical move two months earlier -- but OpenAI's own testing shows detection collapsing once a tenth of the words change.
OpenAI will begin adding an invisible watermark to text ChatGPT and Codex generate for users in the EU, rolling out over the coming weeks to eligible users on every plan tier. The method, called "textGrain" and co-developed with researchers from the University of Pennsylvania and Yale, works by subtly shaping the model's word choices into a pattern a detector can recover using a secret key -- without any visible markup, and without surviving as a per-user identifier. API developers outside the EU can also opt in now; the feature is off by default. The move answers the EU AI Act's transparency rules, in force since August 2, which require labeling AI-generated content.
Anthropic got there first
OpenAI isn't the first mover here: Anthropic announced its own near-identical answer to the same rule in August, watermarking Claude's EU output using Google DeepMind's SynthID Text method (see Anthropic is watermarking Claude's text output, starting in the EU). Both companies landed on the same underlying idea -- steer low-stakes word choices rather than add visible markup -- because it's close to the only approach that satisfies "detectable but invisible" for free-form text at all.
The honest part of OpenAI's own disclosure
The detail worth sitting with is the one OpenAI didn't have to publish: its own tests found that replacing just 10% of a passage's words with synonyms drops detection from roughly 92% to 66%. Short passages, math answers, and translated text are harder to detect in the first place, and OpenAI is explicit that a missing watermark proves nothing about whether a human wrote the text. That's a compliance mechanism doing exactly what a compliance mechanism can do -- satisfy a legal labeling requirement -- and not what a lot of people will assume it does, which is reliably catch AI-written text that someone has bothered to edit. For any organization treating "is this text watermarked" as a meaningful signal about AI usage inside its own walls, OpenAI's own number is the one to remember: a little editing effort erases most of the detection accuracy.