Europe has introduced new regulations requiring all artificial intelligence-generated content to be identifiable. As the rules take effect, companies across the industry are adjusting their systems to comply. While images created by AI can often be traced through metadata, identifying text generated by large language models (LLMs) is more complex. In response, many companies are adopting similar strategies to embed identifiers into their AI outputs. Following Anthropic’s introduction of an invisible watermark in its AI texts this summer, OpenAI has now unveiled its own solution, called textGrain.
The textGrain system works by subtly altering the way an AI model chooses its words. Normally, an AI selects each word based on the probability that it will appear in the current context, with a small amount of randomness that can make outputs feel more natural or unexpected. With textGrain, this randomness is no longer the sole factor. Instead, the model also considers words that are likely to indicate the text was generated by an AI. This deviation in word choice is designed to be imperceptible, ensuring that the text remains coherent and meaningful. OpenAI claims that this method is more refined than its competitors’ and promises only minor changes to the output.
However, the system has its limitations. For accurate identification, the text needs to be sufficiently long—only 80% of 200 tokens can be reliably identified, while the success rate rises to 95% for 400 tokens. Additionally, the detection is more effective with psychological or narrative writing than with technical or mathematical content, which limits the number of possible substitutions. Even minor edits to the text can significantly reduce the effectiveness of the watermark. For example, replacing 10% of the words can reduce detection accuracy by about 30 points, and changing 25% of the words can make detection nearly impossible.
Despite these challenges, OpenAI has decided to deploy textGrain in the coming weeks to meet the requirements of the European Union. Unlike Anthropic, which has rolled out its watermark globally, textGrain will initially be available only to European users of ChatGPT and Codex. This regional rollout will allow OpenAI to assess the real-world impact of the system on a large scale. For developers using the AI models through an API, textGrain is already available upon request, but users must be responsible for applying the watermark themselves. OpenAI has not yet released a public detection tool for textGrain, and only a limited number of researchers and approved organizations will have access to it. In the future, the company may expand access, though no timeline has been announced. The LLMs use tokens to process text, and the number of tokens roughly corresponds to 140 words in French.
European AI Act Drives Adoption of Text Watermarking by Major AI Companies
AI-rewritten from original reportingHow it works
ai-identificationtextgraineu-regulationopenaillm-watermarking



