BitMark watermark was detectable with just one percent of training data marked, study shows

A new study from CISPA shows that watermarks in AI-generated images do not necessarily disappear when those images are used as training data for a new model.

Macro illustration of stacked translucent gelatin sheets with a faint imprint showing through the upper layers, a visual metaphor for watermarks remaining detectable in a model trained on marked data.
Illustration

BitMark watermark was detectable with just one percent of training data marked, study shows

A new study from CISPA shows that watermarks in AI-generated images do not necessarily disappear when those images are used as training data for a new model. But durability varies sharply between methods, and detection works only at the dataset level — not for individual images.

Researchers at the Helmholtz center CISPA have examined a question that is becoming increasingly practical as major laboratories roll out watermarking: what happens to a watermark when the images it sits in are used as training data for a new generative model? The answer, according to the study "Watermark Degradation Across Model Iterations," presented at the ACM IH&MMSec '26 workshop in Florence, is that some marks are removed after a single round of training — while others can be traced even when watermarked images made up as little as one percent of the training data.

How the researchers did it

Michel Meintz of CISPA's SprintML lab and his colleagues tested four watermarking methods — TrustMark, StableSignature, TreeRing, and BitMark — on two different model architectures: diffusion models and autoregressive image models.

The experiment had two parts, Meintz explained in the Helmholtz Association's press release: "First, we used a model to generate images containing watermarks. Then we used these AI-generated images as a dataset to train another model. Our focus was on the training process of this second model. After training, we use the second model to generate new data and check whether the watermark can still be detected or not."

Earlier research on watermarks has primarily dealt with whether they survive manipulation of the image itself — cropping, compression, editing. This study moves the question one step back in the chain: from the image to the model that learns from it.

Some methods die quickly, others hold the line

The results were sharply divided. According to the press release, some methods lose their signal after just one round of training, while others remain detectable considerably longer.

The strongest performance came from BitMark, developed by SprintML itself — that is, the same lab behind the study, something readers should keep in mind when weighing the result. In the Infinity-2B model, the watermark was reliably detectable even when only one percent of the training data was watermarked. For the other three methods, the available sources give no concrete robustness figures; it is stated only that "some" methods lose the signal quickly.

Another important caveat: the researchers did not examine individual images. Instead, they performed a joint statistical analysis of the watermark signals from different image datasets, aggregating weak signals that can be difficult to interpret in a single image. This means the result does not imply that any AI-generated image can be traced back to its training source — only that traces can remain detectable at the dataset level under the studied conditions.

Why this matters

Two motivations lie in the background. The first is model collapse: when companies train on their own synthetic data, quality can degrade over generations. "When companies train on their own synthetic data, it can lead to model collapse," Meintz explained. "The more you train on your own synthetic data, the greater the likelihood that the model's quality degrades." Watermarks that survive training could in principle help filter synthetic data out of training sets.

The second is a practical detection problem across providers: "If companies use different watermarks, one provider cannot easily determine whether an image was synthetically generated," said Meintz, pointing out that companies can only detect their own watermarks and may therefore overlook AI-generated images carrying another provider's mark. The study's finding that marks can survive the generational handoff makes this cross-provider problem more concrete — the traces may exist, but no one is reading them.

Caveats and open questions

The reporting on the study comes from a single source: the Helmholtz Association's press release, republished by Techxplore and Knowridge. Both articles are built on the same institutional source material, so the claims have not been independently verified, and the research paper itself was not available in the sources reviewed.

In addition, the results apply to four watermarking methods and two model architectures under specific conditions. They cannot automatically be generalized to commercial watermarking schemes such as Google's SynthID. There is also no simple rule: a method that works well with one model architecture may be less useful with another, according to the reporting.

The researchers themselves point to two directions for further work: watermarks that remain robust across different systems, and detection using fewer images. Practical use would also be served by compatible detection methods across organizations.

For now, the most cautious message is that retraining does not necessarily erase a watermark — but that relying on hidden marks to trace the provenance of AI-generated material requires testing across multiple model generations first.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.