Back to the blog
AI Regulation

OpenAI Will Watermark ChatGPT and Codex Text, but Only in the EU. Its API Ships With the Switch Off. Claude Ships With It On.

October 7, 2026 6 min read

On Monday, October 5, OpenAI published "Our approach to EU text provenance rules". It says three things.

Starting that day, API customers anywhere can opt in to text watermarking for select models; by default it stays off. Over the coming weeks, eligible text output from ChatGPT and Codex gets an invisible watermark, on all plans, in the European Union only: "We are not making text watermarking a global default at launch." And the detector is not public. Approved researchers and expert organizations can apply for access.

The method is called textGrain. It nudges the model's word choices using a secret key, and a detector holding the same key checks whether a passage follows that pattern more often than chance would allow.

The reason is the EU AI Act. Article 50(2) has applied since August 2, 2026: providers of AI systems that generate text must ensure the outputs are "marked in a machine-readable format and detectable as artificially generated or manipulated." The Commission's guidelines describe a grandfathering rule from the recently adopted AI Omnibus: generative systems already on the market before August 2 have until December 2, 2026 to meet the marking duty. ChatGPT and Codex were on the market long before August, which fits "the coming weeks".

In September I wrote about why "provenance" packs three different claims into one word. This piece is narrower: code, robustness by OpenAI's own numbers, and what the EU-only default means if you ship AI text to European users.

Your repository is mostly out of scope

My first question, as someone who lets agents write code every day: does Codex now leave a watermark in my repo?

The law says no. The Commission's guidelines on Article 50, whose content it approved on July 20, list what falls outside the marking duty, and source code is on the list: anything written in a programming, scripting, markup, query or configuration language meant to be interpreted, compiled or executed by a computer, possibly including the comments that form an integral part of it. SQL, infrastructure as code, YAML and JSON configuration are named. So are short outputs like single words and UI labels, and agent-to-agent traffic that no human sees.

OpenAI's help center agrees on the technical side: "Code is also harder to watermark because there are fewer plausible choices for what comes next than in ordinary prose." The Code of Practice adds a floor for all text: no watermark is required under 200 tokens, which OpenAI puts at about 150 words of English.

OpenAI does not say that Codex skips code. It says "eligible" output. What I would treat as eligible in a Codex session is the prose around the code: a long explanation in the chat, a pull request description, a design note. That is text written for people, not instructions for a compiler. This is my reading, not a ruling.

Two practical consequences. First, there is nothing to see. OpenAI says the watermark adds no hidden characters, no invisible spaces and no unusual punctuation. Diffs, linters and formatters will see nothing, and there is nothing to grep for or strip. Second, unless your team is an approved research or expert organization, nobody on it can check either way.

How robust it is, according to OpenAI

OpenAI is frank about the limits. All of its numbers below are at a target false positive rate of 1%.

On psychology answers, the detector found the watermark in about 80% of 200-token passages and about 95% of 400-token passages. On mathematics, where word choice is less flexible, detection was "substantially lower". In 400-token passages, replacing 10% of the words with synonyms took detection from about 92% to 66%. Replacing 25% took it to 17%. In a separate test across the 24 official EU languages, detection ranged from 69.0% in Spanish down to 42.2% in Romanian, before OpenAI turned up the watermark strength for languages below 60%.

Code is the mathematics case: precise, constrained, often short. The help center adds that substantial paraphrasing or translation can make the mark undetectable. And a 1% false positive target means that, by design, about one passage in a hundred with no watermark at all can come back positive.

The cost to quality looks small. On OpenAI's benchmark table for Astra, which it calls its latest frontier model, watermarked scores move in both directions: DeepSWE v1.1 from 72.80% to 71.68%, Terminal-Bench 4.0 from 53.90% to 56.06%.

So the mark is a signal for an expert holding the key, on long prose nobody edited. OpenAI says it itself: "The absence of a detected watermark does not prove human authorship."

The EU default is a ChatGPT setting, not your compliance

The EU default covers ChatGPT and Codex. It does not cover your product calling the API, even if every one of your users lives in Lyon or Munich. In the API, watermarking stays off until someone turns on "Allow text watermarking" under Organization settings, Data controls, Text provenance, or in a project's settings, and selects the models.

Under Article 50(2), the marking duty sits with the provider of the generative AI system. If you put a product that writes text on the market under your own name, that is very likely you: the guidelines talk about "downstream AI system providers" built on someone else's model. They let you rely on the marking of an upstream model provider, but "without prejudice to the responsibility of the provider of the AI system to demonstrate compliance". A switch nobody flipped is not a marking solution.

Now put the other big vendor next to it. Anthropic marks at the model level, worldwide, on every surface it lists: the API, the Claude apps, Claude Code. Its support page shows text watermarks for Claude Opus 5.5 and Sonnet 5.5, on its own platform and through the cloud partners. It also shows no text watermark yet for older models such as Opus 4.6 or Sonnet 4.6.

I build on Claude, and that table was the useful surprise. The same feature, with the same prompt, produces marked text on a recent Claude model, unmarked text on an older one, and unmarked text on OpenAI's API by default. If you keep a second provider as a fallback, and I think anyone building on a single vendor should, your marking changes every time the router switches.

The guidelines also treat AI-generated summaries as content that needs marking, while translation and grammar fixes count as standard editing and are exempt. And OpenAI says plainly that watermarks "do not replace visible labels".

What I would do on Monday

List every place your product returns generated text to a user in the EU. For each one, write down the provider, the model, the typical length, and the kind of output: summary, draft, translation, code, short label. Long, freely generated prose is what the rule targets. Code, short labels and translations mostly are not.

If any of those paths call OpenAI's API, make an explicit decision on the Text provenance switch, per project, and write it down with the date. Off by default is not a decision.

If you route between providers or models, treat marking as a property of the route, and log which model produced each output. The Code of Practice lists logging as an optional supplement for text, never as a substitute for the mark, but it is the only record you can actually query.

And do not build anything that depends on detection, not a plagiarism check, not a "was this written by AI" filter. You cannot get the detector, and a negative result proves nothing.

The EU-only default shows how OpenAI sees this mark today: something it ships where the law asks, while it learns. Fair enough. For your product, the decision is still yours.

Sources

A project like this one?

I design and deploy products like this. Let's talk.

Let's talk