Imagine asking AI a question with one clear answer. It is ready to say that a meeting will take place on May 12. But the generation system used in the study adds a text watermark so the output can be identified later.
No stamp appears on the screen. The type does not change color. Yet the final answer says the meeting is on May 21.
It is like adding an invisible anti-counterfeit label to a package, only for the labeling machine to change the expiration date while applying it. The package is easier to track, but its contents are harder to trust.
This is an illustration of the generation path described in the paper, not a real user case. In the experiments, all six representative AI text-watermarking methods reduced factual accuracy. Numbers and dates were especially vulnerable.
How does a KGW watermark get into the text?
Think of a large language model as an AI cookie factory. It does not bake the entire box at once. It chooses one ingredient, or the next token, at a time until the passage is complete.
KGW is one method for watermarking text. The factory does not need to rebuild its machinery or print a visible logo on the finished product. It uses a secret key to nudge the choice of each next token.
The cookie factory analogy works like this: the factory quietly favors a set of ingredients known only to itself. A customer cannot taste the marker, but an inspector with the recipe can see an unusual pattern across the whole box. This is a statistical judgment. One greenlisted token alone does not prove that AI generated the text.
That is also where the problem begins. If the correct date receives no boost while a plausible but incorrect date is on the greenlist, the watermark can change the model's final answer.
An illustration of the generation path in the paper: the watermark nudges one token, and the rest of the answer can follow a different path.
One nudged token can pull later facts off course
AI generation is a chain. Each selected token immediately becomes part of the context used to choose the next one.
The researchers identified two main ways watermarking can damage factual accuracy:
Current-token perturbation: To favor particular tokens, a watermarking method can reduce the relative advantage of the correct answer, such as 12, and allow a fluent but wrong alternative, such as 21, to win.
Attention drift caused by the prefix: If the watermark changes meeting to session early in the sentence, that small change becomes part of the model's preceding text. As generation continues, the model may pay less attention to the key evidence in its reference material.
The first change may be small. Later factual differences can accumulate from it.
Factual accuracy fell across all six tested watermarking methods
In Invisible Ink, Visible Lies: How Production Watermarking Causes LLMs to Hallucinate, the researchers ran a controlled comparison.
They gave models including LLaMA 3.1, Mistral and Qwen the same Wikipedia material and questions. Under the same retrieval-augmented generation, or RAG, conditions, they compared answers generated with and without watermarks.
Factual accuracy decreased across all six representative watermarking methods they tested.
The errors were easy to miss. The sentences remained grammatical and coherent, while an important number, date or proper name could be wrong.
The researchers also proposed two interventions, FPTI and FPAI. One limits watermark interference around factual tokens, while the other restores the model's attention to the supporting context. At a detection true-positive rate of 90% and a false-positive rate of 1%, using both interventions reduced watermark-induced factual errors by about 90% compared with watermark-only decoding, although a residual gap remained. Watermarking may be improved, but factual accuracy has to become part of how developers evaluate it.
If watermarks can add errors, verification cannot stop at AI detection
The study leaves us with a useful distinction:
Identifying text as likely AI-generated does not prove that the text is correct.
When an AI answer reads smoothly and sounds certain, you cannot tell from the screen whether an error came from the model, the prompt or a watermark that affected generation.
If the answer will affect a report, purchase, medical decision, legal matter, payment or trip, fluent wording is not evidence. Find the original source and check whether its dates, numbers and names match the answer.
A watermark may help a platform identify AI-generated text, but it cannot verify the facts for the reader. If the marker itself can increase the risk of hallucination, checking the original sources becomes even more important.
AI FactScan organizes the links attached to an answer so you can see whether its claims have sources behind them.
This study does not settle the question for every AI system
The research tested specific open-source models, a controlled RAG setting and a limited set of factual categories. It does not show that every watermark used by commercial AI systems such as ChatGPT or Gemini causes the same effect. It also does not mean every watermarked passage contains an error.
The researchers are not arguing that watermarking should be abandoned. Their point is narrower: watermark evaluation must test whether the method changes factual accuracy, not only whether the mark can be detected.
AI FactScan