Recently, Anthropic announced that all text produced by its models will be watermarked, making your entire interaction not only recognizable as coming from Claude, but also traceable to your identity. According to their press release, the watermarking will survive any copy and paste, and it will be “invisible” to the user.
Two things worth mentioning here.
First, this is a regulatory measure, it’s a EU law that is enforced starting August 2. FWIW, Anthropic did more than the law asked for, but let’s leave this for another time. For now, just understand that this is coming from the top to the bottom, it’s not one AI player going rogue. Everybody will follow suit (Google already did it, even open sourced their SynthId SDK for this in 2024).
Second, the so called watermark is a steganography technique used to hide a message in plain sight, by encrypting token choices. In other words, Anthropic recognizes the text because it choses the next token according to a proprietary algorithm.
So, the output of the model is now permanently “styled” in such a way that it will always be traced back to the model.
Unless the text changes, that is. If you edit the output in a very meaningful way, the token choices are now broken, and the text is “free floating”.
Proprietary Models Are Locking You In, Open Source Models Not So Much
If you still use proprietary models for text generation (docs, blog posts, emails, website copy, etc) keep in mind that the model you used will follow you everywhere. Something you wrote today will still be tracked back to you 5 years from now.
In a (not so) dystopian scenario, by proving that they were part in the generation, AI neolabs can even claim ownership and ask you money for it (even though you already paid for the generation itself). This is not happening yet, to be clear, but nobody says it won’t, either.
There are two ways out of this permanent white surveillance:
- always edit your texts thoroughly – this will break the algorithm and make it unrecognizable
- use open source models which are clearly not watermarking their output
Now, for the first part, the editing. I think it’s worth mentioning that the longer you do this, the more confused AI neolabs will be about who is the text generator, and, in time, it will rend the entire watermarking thing obsolete. No one will rely on something with so many false positives.
As for the second, part, we may be facing a very urgent choice: collectively train models using distributed software, to build clean, open source AI. Everything will be out in the open and we will know the number of params, the training patterns and whether or not is there any watermarking or other stupid surveillance going on.
It won’t be easy, I reckon. But we already did this with money: we started to mine Bitcoin 16 years ago, and look how far we’ve come.
The clock is ticking.
Top comments (0)