产品多源确认79°

OpenAI 将在欧盟为 ChatGPT 和 Codex 文本嵌入隐形水印

精选理由

OpenAI 在欧盟上线 textGrain 水印,400 token 文本检出率约 95%,但改写 25% 就掉到 17%,数据挺实在,做内容的可以看看。

OpenAI 在欧盟为 ChatGPT 和 Codex 输出的文本加入基于 textGrain 方案的隐形水印,通过调整用词模式留下统计信号。在 1% 误报率下,检测器对 400 token 的文本段检出率约 95%,但 200 token 段降至约 80%,数学类答案更低;替换 25% 同义词后检出率从 92% 掉到 17%。该措施对应欧盟 AI 法案第 50 条,所有付费与免费用户将在未来几周陆续启用,全球 API 客户可立即在部分模型上开启。检测工具暂不向公众开放,仅限通过审核的研究者和机构使用。

原文 · rohanpaul_ai

OpenAI will invisibly watermark ChatGPT and Codex text in the EU.

The watermark goes only into text that ChatGPT or Codex writes. So if someone in the EU asks ChatGPT for a paragraph and pastes it into a personal blog, the hidden signal travels with those words, because it lives in the pattern of word choices rather than in any special characters.

If that person rewrites a lot of it, the signal weakens and may disappear. Anything a person writes on their own carries no watermark at all.

Eligible users on every plan get the watermark over the coming weeks, while API customers worldwide can switch it on today for select models.

the public can't use OpenAI's tool for checking whether a piece of text carries the watermark

The rollout answers Article 50 of the EU AI Act, and generative systems already on the market before August 2 must carry machine-readable marks by

Nobody can see it by looking, because the text reads exactly like normal text: no marks, no odd characters, nothing a reader, editor or teacher would notice.

The only way to find it is to run the text through OpenAI's detector, and right now OpenAI keeps that tool to itself plus a small group of researchers and expert organizations who apply and get approved one by one.

The scheme, textGrain, tilts the model's word choices to leave a statistical signal invisible to readers but measurable by a detector.

At a 1% false positive rate, that detector caught about 80% of 200-token psychology passages and about 95% of 400-token ones in OpenAI's tests.

Mathematics answers scored substantially lower, because tight wording leaves fewer interchangeable words to carry the signal.

Editing erodes it too: swapping 10% of words for synonyms in 400-token passages cut detection from about 92% to 66%, and swapping 25% cut it to 17%.

A positive result identifies no user, prompt or owner, while a negative one proves little, since short, edited, translated or rival-generated text escapes detection.