AI Watermark Robustness: Evaluating Methods to Identify AI-Generated Content After Editing or Compression
DOI:
https://doi.org/10.71143/qrzzwv89Abstract
The explosion of Generative Artificial Intelligence (AI) systems that produce photorealistic images, natural-sounding audio, coherent video and fluent text has intensified the demand for dependable mechanisms to trace data back to its source to establish a provenance of synthetic content. One of the most used technical solutions to this requirement is digital watermarking, which consists of adding a slight and imperceptible signal to media produced. In the real world, however, very little content can be presented to a verifier without some alterations: it is often cropped, re-sized, re-encoded, filtered, paraphrased, or deliberately attacked to remove identifying signals. In this paper, we provide a systematic survey of image, video, text, and audio AI watermarking methods, focusing on their resilience to post-generation editing and lossy compression. We categorise approaches into two types of embedding: post-hoc and in-generation and review the threat models and evaluation metrics employed to assess robustness and then compare representative techniques such as hybrids in the frequency domain, learned encoder-decoder-based approaches, diffusion process watermarks, and token-level language model watermarks. We also explore attacks, such as adversarial removal, regeneration and forgery attacks, which are specifically targeting the watermark persistence, and finally collate benchmark results that show a persistent trade-off between robustness and quality/capacity. It ends with open problems including the lack of standard robustness measures, vulnerability to surrogate generative attacks, and the absence of cross-modal or multilingual robustness guarantees, as well as some further directions that are promising, such as semantically based watermarking and certified robustness frameworks.
Downloads
Downloads
Published
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.







