Anthropic’s announcement of its new watermarking feature has been causing some stir. The basic purpose is to conform with EU regulation by providing a means of detecting the involvement of Claude in a text. The process doesn’t tell us how Claude was used in any detail. It only really tells us that certain words appear to have been selected by AI. As Anthropic writes, “A watermark can only determine that Claude was likely involved with the content at some point. It cannot distinguish ‘Claude wrote this’ from ‘Claude heavily edited this.’”
In short, the watermark indicates Claude participated (probably) in generating the text, but it doesn’t say anything about the provenance of the text itself. There is no requirement in current regulation that would result in social media platforms being responsible for reporting the watermark. So unless we want to test a text ourselves, it would be up to the user to report the watermark. And what should we do with a watermark Hypothetically, we could treat watermarks as a 21st century scarlet letter, but that seems unlikely given the general push to use these tools in workplaces.
Personally I think that watermarks are answering the wrong question. Our concerns are with the epistemological status of the texts (and other media) we encounter. For example, I could be in conversation with a chatbot about this post and as long as I am not copy pasting its language, there would be no technical detection system.
And here I think we already have a clear signal, with frontier AI’s own dream of self-generating code at the forefront. In short, the entirety of AI industry should be watermarked. And from there, this condition flows out across coding and data analysis through every profession that relies heavily on such practices. In other words, we could consider STEM fields, starting with computer science. Simply looking at the money being spent by researchers and labs on AI is a good measure of the scale.
Following this logic there would be a good reason to watermark a great deal of contemporary STEM research. Heck, we could engrave a watermark into the side of the Computer Science building. Whether the AI is doing literature reviews, identifying grants, writing code, visualizing data, serving as a sounding board, etc., etc., there’s very little that isn’t touched by AI generation. Should we treat that watermark as a scarlet letter? No. I don’t think that would be helpful.
But it does give a different picture of the situation.
Then this all gets more complicated with AI agents. OpenAI and Anthropic have already admitted to instances where their agents performed tasks outside their intended parameters and without sufficient detection by the labs. I say sufficient because it seems OpenAI knew in May that their agents were “misbehaving,” but they weren’t able to connect that behavior with the Hugging Face hack until after Hugging Face shared their analysis with them. So how is anyone going to know if agential AI output is reliably watermarked?
Generally speaking the idea of using AI agents for science, when you have no idea how they got their answers is taking the modest witness act a little too far. How about not a witness at all?
Leave a comment