DEV Community

Cover image for The End of Undetectable AI Text? Claude’s New Watermark Explained

The End of Undetectable AI Text? Claude’s New Watermark Explained

Sylwia Laskowska on August 11, 2026

For the past few hours, the whole world,  or at least my LinkedIn feed, has been talking about one thing: Anthropic signed the EU AI Act’...
Collapse
 
francistrdev profile image
FrancisTRᴅᴇᴠ (っ◔◡◔)っ • Edited

Thanks for informing us Sylwia!

Sure, the flood of AI slop across social media is annoying. Unfortunately, I have a feeling this watermark isn't going to magically solve that problem either. 😐

To be fair, this is expected and not surprised. It has become an issue and I have receive complaints from other people that they despised AI Slop on this platform (regardless of the moderation I have done).

The best things we can do, is the thing we need to for ourselves, which is start to critical think. Our attention span is so awful to the point where everyone is starting to lose this ability to think for ourselves instead of just giving up and following the herd. Yes, AI is everywhere BUT we need to train ourselves to take control and think to ourselves instead of reaching conclusions and I can't stress this enough.

Sorry Sylwia if this sounds frustrating, but it's something I thought about while typing lol. xD

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Thanks so much, Francis! And exactly, no watermark is going to save us here. Critical thinking is absolutely key.

Honestly, I don't think it's that bad here on DEV yet, but LinkedIn is a nightmare. 😂 And somehow people still like and share all that AI slop, so apparently... they actually enjoy it? Which might be the most terrifying part of the whole thing. 😅

Collapse
 
francistrdev profile image
FrancisTRᴅᴇᴠ (っ◔◡◔)っ

Oh for sure, LinkedIn is a hive for AI slop. It's already toxic enough for people to compare yourself to other people on that platform and it's a no go. Yea sure, I have a LinkedIn, but I rarely post on there.

Honestly, I don't think it's that bad here on DEV yet

It's bad enough where I get reports from @codingwithjiro, @klaudiagrz and other people to the point it's annoying. It's not bad, but it's to a degree where we notice it. Sure, people use it to a good degree for Grammar, but people ruin the fun for that by claiming that, where they didn't. Oh yea, for comments, still have no idea why we see AI-Generated Comments. I can understand posts and all, but for comments is crazy. I understand professionalism, but there are many ways to do this. Was wondering if you knew why this maybe the case?

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska

Yes, AI-generated comments... It's CRAZY!!! IDK, maybe it's a kind of self-promotion? I even understand AI-polished comment, for clarity. But some of them are pure AI slop 😅
Anyway, I'm going to sleep now, I'm not sure what I'm doing here on DEV at 1 AM 😂

Thread Thread
 
klaudiagrz profile image
Klaudia Grzondziel

OMG, YES! 💯 I understand running a spellcheck over the article to correct typos and so on... but for the comments... come on, I want to hear your voice, your style, your way of expression, not another generic slop 😬 We all communicate in different ways and styles, and that's what's interesting about exchanging opinions on such platforms like DEV.

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska

I can even understand running a spellcheck on a comment. But some comments are literally just generic AI-generated summaries of the article with no original thought added at all. And at that point I don't understand why someone even bothers posting them 😅

Collapse
 
anchildress1 profile image
Ashley Childress

Thanks for the writeup. Tbh, I'm glad somebody is finally enforcing this AI content thing. I wrote about it a while back when I got upset with the "ban everything AI" wave that's still hanging around in places. My biggest concern with this is that your distinction

watermark detected ≠ AI wrote everything
no watermark detected ≠ human wrote everything

gets collapsed into something resembling the AI version of hellfire and pitchforks...

Great article, nonetheless!

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Thanks so much, Ashley! I'm starting to notice this panic too, especially among non-technical people: this idea that anything involving AI is automatically bad.

But people are going to use AI anyway. And just like @francistrdev said, I think we simply need to get better at critical thinking instead of treating everything as either "AI = evil" or "AI = good." 😅

Collapse
 
sushyam_nagallapati profile image
Sushyam Nagallapati

@anchildress1 Absolutely agree. The nuance gets lost way too easily. Once people see a positive detection, nobody stays around to ask whether it was a full draft or just a quick grammar pass before the pitchforks come out.

Collapse
 
gramli profile image
Daniel Balcarek

Interesting, thanks for the information, Sylwia. I’m really curious how this will work in practice, especially with code.

At least in our case, we use AI to generate code that has to fit into an existing codebase: naming conventions, patterns, formatting ... So I’m really curious how they want to put a watermark into that and how much of it will survive after the code is adapted, formatted, or refactored.

It looks like I have the wrong friends on LinkedIn though, because my feed definitely didn’t blow up with this. I still mostly see recruiter posts with pictures of them working from the beach. 😂😂

PS: I just had to add this, it perfectly fits the whole EU situation (and it doesn’t even have to be about watermarking). 😂

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Hahaha, maybe we'll start getting variables like claudeResponse = await claudeFetchData() just to make sure the watermark survives. 🤣

BTW, I'm genuinely jealous of your recruiters working from the beach! My LinkedIn feed is basically permanent drama, especially Polish LinkedIn. 😂 A very popular genre of post there is basically: "EVERYONE ELSE is publishing AI-generated garbage, but NOT ME. I’m honest, I do everything properly, and I’m better than that."

Collapse
 
gramli profile image
Daniel Balcarek

The beach photos are nice, but all those recruiter mottos and “smart sentences”… I’m honestly getting tired of them, but even so, your feed sounds much more irritating. 😂😅

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska

Oh yes, mine is absolutely terrible. 😂 It's basically: "Everyone on LinkedIn is doing it WRONG! Someone dared to ask a question at the end of their post, probably just for engagement! How dare they! Only I do LinkedIn properly and honestly!" And then underneath: 1,000 likes and hundreds of comments saying, "YES, this annoys me too!"

Meanwhile, LinkedIn itself... if I post something relatively smart, I get maybe 10–20 likes. But if I post a photo of myself making a stupid face? 200+. xDDDDDD

So apparently the algorithm has spoken. 😂

Thread Thread
 
gramli profile image
Daniel Balcarek

😂😂😂 And if that photo was from the beach, it would probably be 2,000+. 😂😂

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska

That's the next step, once I decide it's time to achieve international fame. 😂😂😂

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Ah, Sylwia! Your articles always seem to show up exactly where I need them. I was actually wondering about this watermark earlier tonight.

And what Claude says — and what you point out — pretty much confirms my view: this is a band-aid on a wooden leg.

Written by AI? Written with AI assistance? Corrected by AI? Not AI at all? Or actually written by AI but simply not detected?

In other words, complete nonsense on a scale matching the sheer foolishness of people who think they can know everything about everything and solve every problem with a binary answer.

It reminds me, for instance, of the proposed ban on social media for under-15s, or France’s tax supposedly aimed at blocking Temu, which ultimately did little more than hurt jobs in France — and, incidentally, make Bercy and its brilliant minds look rather ridiculous.

The more complicated the problem, the more absurd it becomes to pretend there’s a simple yes/no solution.

By the way, I wrote this comment in French and had ChatGPT translate it into English. So… how exactly should we classify it? 😉

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Exactly, Pascal! And for many people, this is simply a way of leveling the playing field. We're not writing novels here. 😅 Someone translates their text with AI (I sometimes write comments in Polish and translate them too, simply because it's faster), someone else uses it to fix grammar and typos, and so on.

And Anthropic explicitly says that this kind of AI-assisted text can get the watermark too. That's it. So what exactly would detecting that watermark prove?

That's also why I suspect these detection tools won't be universally available to everyone. Otherwise, I can already imagine the absolute paranoia we'd end up with. 😂

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

I suspect these detection tools won't be universally available to everyone

I hope so, but unfortunately, I have absolute faith in our institutions: if faced with a choice between an intelligent decision and others that are more questionable, they will invariably choose one of the worst options. They’re basically applying the Dilbert principle...

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska

Hahaha, I'm dying. 😂 I thought we only had decision-makers like that in Poland, but apparently this is a truly global phenomenon.

Collapse
 
klaudiagrz profile image
Klaudia Grzondziel

Interesting article, Sylwia! Thank you. I didn't know about Anthropic's watermark.

I'm really curious what the future will bring... I noticed that I myself tend to quit reading the content as soon as my brain realises it is generated. Introducing a watermark can bring a new trend of people boycotting AI slop by simply not reading it. So not only people like me, who easily get AI fatigue, but also people who reject content simply because it is generated with AI. This may, in turn, bring us back to the times when people put more heart into writing, not just prompting AI to do this for them. Or maybe it's just my wishful thinking 😅

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Maybe! 😄 Although knowing how these things usually go, I suspect the people we'd most want to protect ourselves from will be the first ones to figure out how to bypass the watermark.

Collapse
 
wiseai profile image
Mahmoud Harmouch

Good read! But no matter how hard they try to watermark anything, their efforts are ultimately pointless. Any watermarking algorithm can potentially be reversed, even when it relies on statistical techniques, such as systematically distributing the sampling of certain words.

In fact, I’ve built a humble bumble lil tool that lets you remove watermarks from outputs generated by any LLM provider. I called it CUM: Claude Unmarking Machine.

GitHub logo wiseaidotdev / cum

💦 Claude Unmarking Machine.

cum-rs logo

CUM

Crates.io Docs.rs npm PyPI CI Clippy MIT License MSRV

Claude Unmarking Machine: a multilanguage Rust crate that removes AI-provider watermarks from text, images, and documents Works regardless of provider (Claude, OpenAI, Gemini, Grok, open-LLM) All processing is 100% local: no data leaves your machine.

crab dancing

The cum binary, cheerfully evicting zero-width gremlins from your prose.

🤔 What is happening here, exactly?

So you copy-pasted some text from an AI. Totally normal. You are not doing anything illegal. Probably.

The bad news: every major LLM provider stuffs your output full of invisible Unicode graffiti so they can identify their own generation later. It is like spray-painting "CLAUDE WAS HERE" on every wall, except the paint is literally invisible and you cannot see it without specialized equipment.

The good news: we have the specialized equipment. And it is written in Rust, so it is blazingly fast.

Layer What lurks in the shadows What we do about it
A: Unicode


Try It Out

Try It

Some watermarked text you can try to unwatermark:

CUM Examples.

Have fun!

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Okay, everything sounds great, but I would maybe reconsider the name 😂😂😂

Collapse
 
wiseai profile image
Mahmoud Harmouch

Hiya (´• ω •)ノ!

Sowwi for the late response. I'm only online on weekends these days, unfortunately.

Actually, if you think about it, CUM is a very corporate-friendly term. From my experience in my previous onsite corporate role, I remember the office being built around pure "cum" policies. It was one of the most toxic environments I've ever experienced as a software engineer. Luckily, though, I resigned 2 months later, not just because of that, but because of many other factors as well (e.g., extremely low pay, working onsite from Monday till Saturday, 8 AM to 6 PM).

Software engineering should feel much more like playing than some form of insane peak slavery.

Beyond that, we're just having fun here on Dev.

Hope you enjoy reading my articles!

Collapse
 
learn2027 profile image
meow.hair

Hi Sylwia, fantastic breakdown! You perfectly cut through the hype and myths surrounding the EU AI Act and Anthropic's watermark.

Your point about the "Code Watermarking" challenge is the most critical part of this discussion. In natural language, the latent space for synonyms is vast, but in code, syntax is rigid. If we add tools like Prettier, ESLint, or minifiers into the CI/CD pipeline, the statistical bias could easily be flattened out or stripped entirely.

I'm really curious to see Anthropic's technical docs on how they plan to solve the code watermarking problem without breaking linters or refactoring workflows.

Do you have any insight into whether Anthropic has discussed a multi-stage watermarking approach that survives code minification? For instance, minifiers rename variables and remove whitespace—so wouldn't it be more robust to embed the watermark in the AST (Abstract Syntax Tree) structure rather than the raw text? This seems like the real engineering challenge: any watermark that relies on surface-level syntax might be stripped after running Prettier or ESLint.

Thanks for another insightful read! Wishing you continued success and more great technical deep dives!
🗻🌊

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Thanks so much for the kind words! 😊 And unfortunately, I can't answer your question about code because... I simply don't know! 😅

Anthropic hasn't published the technical details yet and says the documentation is still coming, so for now, we have to wait. I'm really curious about this part too, especially how (or whether) the watermark could survive things like formatting, refactoring, and minification.

So... Anthropic, we're waiting for those docs! 😄

Collapse
 
kenwalger profile image
Ken W Alger

Great article, @sylwia-lask, thanks for putting it together.

I find the whole AI-generated content particularly interesting from a job candidate standpoint. Job descriptions for the roles I'm looking for frequently include "must be knowledgeable, familiar with, and actively use AI tools for content and code generation." That's fine. At the same time, they don't want to see anything actually generated by AI. It all must be human-generated. Blog posts, resumes, cover letters, repository code, documentation, etc., are all expected to have "never been touched by AI," but at the same time we need to be well-versed in using it.

I fully understand and can appreciate the concept of "AI Slop." But, in my opinion, there needs to be some happy middle ground somewhere: "This was a human's idea, guided by a human, integrates a human's experience, etc., and was polished in some fashion by AI" that becomes acceptable to folks.

Do I think AI is a silver-bullet solution for content generation? No. In my experience, strictly AI-generated content loses something that humans bring to the table. When writing about technical things, is having an AI available to double-check that your spelling is correct, arguments are sound, and structure makes sense helpful? I think there are cases when it is, yes.

We're just in somewhat of a strange period at the moment, I think, with AI, in which lots of folks want to use it, but apparently no one wants to admit they used it or accept the outcome of it. And corporately, we're seeing it being used more and more, either AI -generated content itself or agentic use cases, so it doesn't appear to be going away. The "AI Slop" issue will be around for a while, I think. Regardless of how many watermarks are added. Someone out there will undoubtedly create a "watermark stripper".

Maybe it's just me and my experience here, and sorry for the rambling.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Hey Ken, I actually really like this way of thinking about it!

It reminds me a little bit of Hollywood, where everyone gets plastic surgery, but hardly anyone wants to admit it because then people will point fingers at them. 😂 Better to insist that no, no, it's all genetics!

And I think it really depends on what you're creating. I'm not looking for a job at the moment, but I do write a lot of different things, so that's where I can speak from experience.

For example, with technical articles or comments, I often use AI to fix my grammar, or I simply write something in Polish and ask it to translate it into English. These are technical texts, often mixed with a bit of essay/opinion writing, and usually have a relatively short lifespan. Their purpose is to communicate an idea, so I don't see a reason to spend hours manually polishing every sentence just to be able to say that AI never touched it.

But when I'm writing fiction, that's a completely different craft for me. I don't let the model change even a comma. I can read the same sentence ten times and obsess over every word, because the writing itself is part of what I'm creating.

And there's another aspect to this, especially here on DEV. There are people who are absolutely brilliant technically, but they're not necessarily writers, and many of them aren't native English speakers either. I prefer that they run their text through AI and make it easier for me to understand their ideas, rather than struggle to write everything completely on their own and end up with something I can't understand.

So yes, I think we desperately need that middle ground you described. AI assistance doesn't automatically make something "AI slop," just like avoiding AI doesn't automatically make something good. 😅

Collapse
 
edmundsparrow profile image
Ekong Ikpe

Watermark stripper 🤣 I can't imagine any further 😅

Collapse
 
newadventuresinit profile image
Dirk Mattig

Thanks for bringing up this important topic, Sylwia!
At this point in time I do not understand what the intended purpose of this technology is.
For example, if it cannot tell apart generated content from generated translations then what exactly does this mark indicate? It seems to carry little information but instead a massive potential for misunderstandings and misuse.
Furthermore, if it is really some kind of "watermark" then rather sooner than later tools will exist to remove them. It would be a first in IT that a protection mechanism survives longer than a few weeks once it is out.
It will be interesting to see how the labs will react in case a widespread misuse of this technology will drive customers away from their models to either competitors or open-source models.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Exactly, Dirk! The more I discuss this with people in the comments, the more I get the impression that the main purpose of all this is simply to be compliant with EU law 😅

Collapse
 
sarahpan profile image
Sarah Pan

From the positive perspective, the watermark could help universities to check whether students cheating or violating academic integrity policies by using AI. They may not be the definitive proof, but they could provide an additional signal.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Exactly! Or for detecting deepfakes in general. That's why I have a feeling access to these detectors might be limited to institutions that actually need them, like universities and similar organizations.

In Poland, for example, universities already use an anti-plagiarism system for academic work. And even that can be bypassed relatively easily if someone is good enough at rewriting things. 😅

So I can imagine watermark detection being treated as an additional signal rather than definitive proof. But we'll see how they actually implement access to it. At this point, nothing would surprise me. 😂

Collapse
 
buildbasekit profile image
buildbasekit

So we finally got AI watermarks.

Now we just need a watermark for the human who approved the PR without reading it.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

hahahaha will it also have false positives? 🤣

Collapse
 
buildbasekit profile image
buildbasekit

😂😂 probably. Imagine getting flagged for approving the PR and for not reading it.

Collapse
 
sushyam_nagallapati profile image
Sushyam Nagallapati

@sylwia-lask The point about code watermarking is spot on. Between linters, auto-formatting, refactoring, and strict team style guides, I really don't see how a statistical token bias survives a standard build pipeline. Text is flexible, but code context is way tighter. Will be really interested to see their technical breakdown once it drops.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

You know what? They actually published an update about this! And they basically admit that code is going to be a problem.

It might be possible to embed a detectable watermark when the model generates a really large amount of code, but often there simply won't be enough freedom to do it. Sometimes there's just one correct solution, so the model can't freely choose between different tokens just to introduce a statistical signal.

Collapse
 
sushyam_nagallapati profile image
Sushyam Nagallapati

Thanks for sharing that update!

It makes total sense, code syntax is way too rigid for a statistical signal to survive without sacrificing accuracy or readability. It'll be interesting to see if watermarking ultimately remains practical only for natural language prose while code requires entirely different detection paradigms.

Collapse
 
sachin_krrajput profile image
Sachin Kr. Rajput

@mudassirkhan19 asked about detection confidence thresholds near the end of this thread and nobody picked it up, so — it is the question with an actual answer, and the answer is that a threshold cannot be chosen from the detector's accuracy at all. It needs the base rate and the cost ratio, and neither is a property of the watermark.

Take a generous detector: 99% true positive rate, 1% false positive rate. Run it over 10,000 essays.

30% are AI-assisted -> 2970 true, 70 false -> 97.7% of flags correct
5% are AI-assisted -> 495 true, 95 false -> 83.9% correct, 1 in 6 flags wrong
1% are AI-assisted -> 99 true, 99 false -> 50.0% correct, a coin flip

Same detector, untouched. Only the prevalence moved. In the course where almost nobody cheats — the one where an accusation is most devastating — half the accusations are wrong.

And the threshold question needs the cost ratio explicitly. If falsely accusing a student is 100x worse than missing one, you need a likelihood ratio above 100, so FPR below 0.0099 — and pushing specificity that far collapses sensitivity, which is the trade nobody writing policy will state out loud.

The detail from your SynthID update makes this sharper: a full translation gets watermarked while proofreading may not. So the positive class is "a model chose these tokens," not "this person cheated." The prevalence of actual misconduct among watermark-positives is lower again, which pushes precision down further.

Which is why "probably watermarked" has no defensible action attached to it until someone names the cost ratio. Everything else is picking a number off a ROC curve and hoping.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Thank you so much for this comment. This is a fantastic addition to the article, and I think it shows really well why a detection score alone tells us much less than people tend to assume.

And from my side, I think it's interesting to look at how differently platforms approach this in practice. Medium, for example, requires disclosure for AI-generated text, including AI assistance beyond things like grammar or spell checking. HackerNoon, on the other hand, actively offers automatic translations of articles into multiple languages. So even platforms themselves clearly don't agree on where exactly the meaningful boundary is.

I also have a friend who works in social media and naturally writes in these very polished, perfectly rounded sentences. ZeroGPT regularly tells her that her completely human-written text is something like 60% AI-generated 😂 Not “polished by AI.” Generated.

And then there's what might ultimately make this whole debate even more absurd: we're going to get a million watermark removers. Some are already appearing. So people who actually want to hide AI use will simply remove the watermark, while honest users may end up being the easiest ones to detect.

Which brings me back to your point: turning any of these signals into an accusation or an automated policy decision is a much bigger leap than simply detecting that a model probably generated some tokens.

Collapse
 
sachin_krrajput profile image
Sachin Kr. Rajput

Your last paragraph makes my arithmetic look optimistic, and I think it is worth putting a number on because it inverts the sign of the whole thing.

I treated the base rate as fixed. If motivated misuse strips the watermark and honest AI-assisted users do not, the base rate is not fixed — it is adversarially selected. Say 10,000 texts: 500 deliberate misuse, 2,000 honest AI-assist (translation, grammar), 7,500 pure human. Same 99%/1% detector:

0% strip -> 2550 flagged, 19.4% are misuse (1 in 5)
50% strip -> 2302 flagged, 10.7% (1 in 9)
80% strip -> 2154 flagged, 4.6% (1 in 22)
95% strip -> 2080 flagged, 1.2% (1 in 84)

The flag count barely moves. What changes is who is in it. At 80% stripping, twenty-one of every twenty-two flags is an honest user, because the detector is now measuring "did not bother to evade" rather than "cheated."

Your friend's ZeroGPT result is the other half, and it is the part that bothers me more. A 1% false positive rate is a population average. It is not her rate. If polished, formal, or non-native-English prose reads as generated, then the errors are not independent draws — they concentrate on the same writers every time. So she does not experience a 1% error rate. She experiences being accused repeatedly, and there is no number of clean submissions that fixes it, because the thing being detected is her style.

A detector whose false positives correlate with the author is not a noisy detector. It is a biased one, and averaging over the population is exactly what hides that.

Collapse
 
wrobeltomasz profile image
Tomasz

I suspect that if you use synonyms for at least 40% of the words, you can blur the watermark. What do you think?

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Yes, very possible, but I definitely wouldn't bet on that 40%. 😄 Maybe changing 20% of the words in the right places would be enough, or maybe even 40% wouldn't be.

BTW, there was actually an update on this today (although I'm not sure I have the energy to write another article about it... 😂), and it turns out Anthropic is indeed using SynthID.

And here's an interesting detail: if Claude only does proofreading, basically fixing typos and small mistakes, the watermark may not be detectable because Claude isn't generating enough new text. But if Claude translates the entire text, the translation can be watermarked because Claude is choosing all the new tokens, even though the actual content and ideas obviously weren't generated by AI.

Collapse
 
wrobeltomasz profile image
Tomasz

I can see for myself pratice that even though the model is supposed to give short answers, model still writes too much. Or is the watermark being forced in this way? 🤨

Thread Thread
 
sylwia-lask profile image
Sylwia Laskowska

I think models just naturally love to talk too much. 😂 They were already doing that long before watermarks.

Thread Thread
 
wrobeltomasz profile image
Tomasz

He especially likes to sneak comments into the code, which is perfect for annotating code generated by AI models.

Collapse
 
publiflow profile image
PubliFlow

Solid ML write-up. For production ML systems, I'd add that implementing proper experiment tracking (MLflow, Weights & Biases) from the start is invaluable — you'll thank yourself when you need to reproduce results months later.

Collapse
 
mortogn profile image
Amorto Goon

Never enjoyed reading AI generated content. If I wanted to learn something from AI, I could've just prompted it myself. At least with this, if it works as intended, will solve the headache most people face with AI generated content like myself.

I didn't know they are implementing it. Thank you.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

I don't think it will be quite that simple. 😅 My guess is that the people we'd most want to stop from mass-generating this kind of content will also be the first ones to figure out how to bypass the watermark.

Collapse
 
sanskari_patrick07 profile image
Prateik Lohani

Personally I think watermarking code in particular seems kinda dumb and I doubt they'll train the models to do so.

With general content at least there's the concern of public opinion and how it can be shaped through models but code is objective logic. There is one use case that could come into picture which is flagging PRs based on how much it's AI generated but other than that I can't think of any.

IDK maybe I'm not thinking hard enough about this.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Yeah, code is definitely a tricky case. 😅 Anthropic actually updated their docs yesterday and basically confirmed that watermarking code is difficult for exactly this reason.

There simply aren't enough arbitrary choices in many pieces of code to reliably embed a statistical signal. If the model generates a really large amount of code, then MAYBE there will be enough opportunities to introduce a detectable watermark, but for smaller snippets it can be very difficult.

So I'm also still wondering how useful code watermarking will actually be in practice. :D

Collapse
 
mk023 profile image
Marco

I think there is an interesting distinction here between provenance and detection. 🔍 A watermark can be useful as a provenance signal, especially for media like images, audio and video, but I don't think it should be treated as proof that "AI wrote this".

The code example makes this even more interesting. 💻 If AI generates code, a human reviews it, refactors it, tests it and commits it, what exactly would an AI watermark tell us? For software, I'd argue that code provenance, review history, testing and supply-chain controls are much more useful signals than simply knowing that an AI model was involved. 🔐

I'm curious how others see this: where do you think AI watermarking provides real value, and where does it become mostly noise? 🤔

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Hahaha, exactly! And that raises another question: what does knowing that a piece of code was generated by an AI model actually give us, especially when pretty much everyone is using AI for coding anyway? What are we supposed to do with that information? 😅

I think watermarking has real value in areas like academic work or deepfakes (although knowing how these things usually go, people creating deepfakes will probably be the first ones to figure out how to bypass it 😂).

But in everyday use? I don't think it matters much. Do I really care whether someone wrote an email themselves or used AI to help them? Sometimes I'm actually happier if AI helped, at least the thoughts might be easier to understand.

Collapse
 
edmundsparrow profile image
Ekong Ikpe • Edited

I think we have to accept that we're all guilty of technology — users and creators alike. 😅

Technology has been making the world lazier for as long as I've been conscious of it. That was actually my mindset about technology long before I ever touched a PC. But I've never believed technology can make the mind lazy in quite the same way.

We can automate more and more of what humans have to do. We can even outsource parts of writing, coding and research. But we can't outsource the fact that a human still has to think, have an idea, question something, or decide what they actually want.

And this is where AI gets particularly interesting to me. At the height of it, I've come to see pattern matching as statistics, and statistics as a function of abstraction. That's part of what makes these assistants feel so powerful. They can recognize and connect patterns across enormous amounts of information.

But abstraction without strong territorial knowledge has limits. The model can produce a beautifully coherent answer while missing the actual territory, especially when the context is ambiguous or the expression is weak. That's where hallucination becomes fascinating rather than simply "AI being stupid."

Humans do something similar too: we pattern-match, abstract and fill gaps from experience. The difference is that our territorial knowledge can sometimes tell us, wait, something is wrong here.

And I'm sure some people will disagree with that — just like some people will still be arguing in 2030 that using AI assistants means you're no longer original. 😂

The world will probably keep getting lazier. I'm not convinced human thinking will disappear with it . Happy weekend ✌️

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Exactly! Critical thinking > AI detection, always. 😄

And when it comes to writing text or code, there is writing... and there is writing. We're often not talking about novelists or poets here, but people from completely different fields who simply want to communicate an idea clearly.

In that case, AI is just a tool. Not that different from how an IDE became a tool for programmers years ago. It helps you express or build something, but you still need to know what you're trying to say or create in the first place.

And happy weekend for you too :)

Collapse
 
edmundsparrow profile image
Ekong Ikpe

✌️

Collapse
 
debashish_ghosal profile image
Debashish Ghosal

Sylwia, this is exactly the post my feed needed this week. Two things I appreciate most. First, the way you nailed the two one-way implications: "watermark detected ≠ AI wrote everything" and "no watermark detected ≠ human wrote everything." Almost every take out there collapses that distinction into "AI detector, so what are you hiding?", and you kept it precise. Second, the code question: in natural language there are dozens of equivalent ways to say the same thing, but const result = await fetchData() leaves almost no token space for a statistical signal — and you were honest that Anthropic hasn't published enough detail to know how they solve it. That honesty matters. My suggestion while we wait on the docs: run your own survivability test. Generate a few hundred words via Claude, then put it through realistic transformations — translate to another language and back, heavy paraphrase, one honest edit pass. Whatever survives that gauntlet tells you empirically whether this watermark actually moves the needle on slop detection. Cheap to set up, and you'd have a great follow-up post.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Thank you so much, Debashish! 😄 And honestly, if you like the survivability test idea, please feel absolutely free to steal it for your own post! My fall has somehow turned completely insane, so realistically, I probably wouldn't get around to testing this properly until November 😂 I'd be very curious to read the results much sooner than that!

Collapse
 
heinrichneb profile image
Heinrich Neb

Thank you for the update in the comments about translation vs. proofreading — I think that one detail is doing far more work than it first appears, and I'd genuinely like your take on where it lands.

If proofreading rarely produces a detectable watermark (too few new tokens) but a full translation reliably does (the model picks every token), then the mark correlates with which language you thought in — not with how much help you took.

Concretely: two people write the same article with the same amount of AI involvement. The native English speaker asks Claude to fix grammar → probably no detectable mark. You write it in Polish and have Claude translate → marked. Same ideas, same authorship, same editorial control, different flag.

That isn't a flaw in the watermark; it does exactly what it says it does. But if any platform ever turns that flag into ranking or reach, the cost lands hardest on people writing in a second language — who are also the people the tool helps most. And they'd be paying it for the part of the work that involved no AI thinking at all.

You mentioned yourself that you sometimes write comments in Polish and translate them. Does that thought bother you as much as it bothers me, or do you see a reason it might not play out that way?

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

You’re 100% right, and I have very mixed feelings about this too. Yes, I sometimes write comments in Polish and then translate them. Personally, I don't think the reader loses anything because of that. What ultimately matters is the content. You can usually tell whether a comment contains an actual thought, argument, or insight, or whether it's just some meaningless paraphrase generated by a model.

And my view actually goes further than translation. If someone uses an LLM while writing an article, not just to translate it, but, for example, has a conversation with it, reaches some genuinely interesting conclusions through that conversation, and then lets the LLM help put those conclusions into words because writing isn't their strongest skill, I'm completely fine with that. The intellectual contribution is still there. The tool helped express it.

But I can also see that for many people, especially outside IT, this distinction is much more binary: an LLM was used or it wasn't. And I think that's a pretty bad way to judge authorship or the value of a piece of writing.

And then there's another layer of absurdity: we're probably about to see plenty of tools designed specifically to remove these watermarks. Some are already appearing. So anyone who actually cares about hiding AI involvement will likely be able to remove the signal anyway, while people using LLMs transparently for perfectly legitimate things like translation may be the ones who get flagged.

Collapse
 
heinrichneb profile image
Heinrich Neb

Your last paragraph is the one that stays with me, because it describes a
pattern with a long history: the mark survives only on the people who weren't
hiding anything. DRM went exactly this way — the friction landed on paying
customers while pirates shipped clean copies. If watermark removers become a
commodity, the detectable population is precisely the transparent one.

On the detector side I can offer one number from our own kitchen, because we
walked into this exact trap last week in a completely different domain. A
quality benchmark of ours scored 92.3% on 17 hand-written test cases — and 5%
when we finally ran it against 499 real records. Nothing was wrong with the
math; the test data just looked nothing like reality. Every public AI detector
advertises exactly that kind of lab number, which makes me read their
percentages with curiosity rather than confidence.

A question for your follow-up, once the technical docs land: would you consider
testing the boundary empirically rather than waiting for it to be documented?
Same text, increasing intervention — typo fixes, grammar pass, paraphrase,
full translation — and check at which step the watermark becomes detectable.
Anthropic mentioned researcher access to detection; you seem like exactly the
right person to claim one of those seats. That one chart would answer the
proofreading-vs-translation question for every second-language writer on this
platform, and as far as I can tell nobody has published it yet. I'd genuinely
love to read that article.

Collapse
 
473185670 profile image
CBT Tools

The distinction that matters most isn't in the article's two-way split — it's one level up: a watermark tells you who wrote the text, never whether the text is true.

I shipped an AI-generated backtest summary to 234 readers across three platforms. It claimed "GOLDILOCKS +1.2% vs CONTRACTION -2.1%" — fluent, confident, well-structured, matched my priors so it read as authoritative. A text watermark would have correctly flagged it as AI-assisted, and that flag would have told readers exactly nothing useful. The dangerous property wasn't "a model wrote these words" (which the watermark catches) — the words were fine. The dangerous property was "this confident claim is false," which no watermark catches. The real event study later showed p=0.643 with the signal backwards at all four horizons.

Your distinction generalizes past provenance:

  • watermark detected != the content is false
  • no watermark detected != the content is true

For AI slop on LinkedIn, provenance detection is the whole game — "did a model write this empty engagement-bait?" is the question. For technical claims with money behind them, provenance is almost irrelevant and truth is the only question, and truth is the thing no statistical token bias can encode.

The code-watermarking subthread is where this gets sharp for me. Even if Anthropic perfectly watermarks a generated function, the watermark says a model wrote it. It says nothing about whether the function is correct, whether the model's "I ran the tests and they passed" claim about it is true, or whether the test suite itself was AI-generated and flattering. Marco's instinct above — code review ignores provenance and reads behavior — is right, and it's right precisely because for code, truth-checking (does it do what it claims) already exists and provenance-checking (who typed it) adds little. The watermark's real value is supply-chain attribution, not correctness.

The critical-thinking thread is the closest, but I'd split it: "is this AI slop?" (provenance, watermark helps) is a different and easier skill than "is this confident-looking claim actually true?" (truth, watermark is silent). Only the second one has money in it, and it's the one my backtest summary passed for three platforms and 234 readers until an event study the agent couldn't write to overrode it.

Open source: real_backtest.py — the detector that actually worked.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Exactly! And even the EU regulation doesn't say that you're not allowed to generate text with AI. You can, even entirely if that's what you want. The important part is that someone properly reviews it, takes editorial responsibility, and is ultimately willing to put their name behind it.

And when it comes to LinkedIn... honestly, I think people just need to start thinking for themselves. 😂 No watermark is going to save us there.

Collapse
 
simouun profile image
Sami • Edited

This article makes me wonder about the open-weights ecosystem (like running Llama or Qwen locally via Ollama). Since everything runs locally on the user's machine, enforcing a watermark seems really difficult. Wouldn't local models simply become the de facto alternative for developers needing full control over their outputs? I’m very curious to see how this side of the AI world fits into the picture.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Absolutely! A locally running model would definitely solve the watermark problem. 😄

Collapse
 
codingwithjiro profile image
Elmar Chavez

I strongly feel like the watermarking feature would be a sophisticated process. How do they even tag a simple AI generated paragraph or code. There will also be errors in judgment too, that's for sure. If a detector is live, it should be used by the public too. The internet is filled with AI slop nowadays that having an automatic detector somewhat becomes a necessity.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

You're right, Elmar! The longer I think about it, the more I feel like the main goal here was simply to comply with the EU requirements, and that's basically it. 😅

Collapse
 
talha_ramzan_3878156fea8c profile image
Talha Ramzan

The code watermarking question is the one I'd want answered most too.
Natural language has huge redundancy to hide a statistical bias in,
code doesn't, variable naming and formatting are about the only
degrees of freedom left once you're past a certain point of
correctness. If the bias has to concentrate into fewer decision
points, it seems like it'd either be more fragile (one refactor kills
it) or more detectable (if it's leaning hard on the few choices
available).

The "watermark detected ≠ AI wrote everything" framing is the part
worth repeating everywhere this topic comes up, most public discussion
collapses that into a binary the moment it leaves a technical audience,
and that's where the real-world false-accusation problems come from,
not the technology itself.

Appreciate that you separated what's confirmed from what's an educated
guess, most coverage of this blurred that line entirely.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Exactly,so many unknowns! Anthropic has now released more detailed docs, and my guesses turned out to be pretty much correct: SynthID for text, while code is going to be very difficult to watermark reliably.

Which does make me wonder… was this whole circus mainly about being able to say “yep, we comply with the EU rules now”? 😂

Collapse
 
syedahmershah profile image
Syed Ahmer Shah

I was actually about to write about this this week 😅. Really interesting development, but I still think humans will always find ways to remove, alter, or work around the watermark. The real challenge is making provenance useful without treating it as absolute proof.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Hahaha, you didn't miss out on much by not writing it! 😅 Despite getting lots of clicks and engagement, my article seems to have been heavily downvoted and has practically disappeared from the feed.

So maybe someone really doesn't want us to know the truth about the watermark. 😂

And yes, I completely agree. The first people to figure out how to bypass these watermarks will probably be exactly the people we wanted the watermarks to protect us from in the first place. 😅

Collapse
 
kenielzep97 profile image
Self-Correcting Systems

the code question at the end is the best thing in here and i dont think it has a comfortable
answer, so let me try to make the shape of it worse.

statistical token watermarking needs entropy to work. you can only bias a choice where a choice
exists. so the real question isnt whether code has any entropy, its where the entropy actually
sits, and in code it sits almost entirely in identifier names, formatting, ordering of independent
statements, choosing between equivalent constructs, and comment text. everything else is pinned.
once the model commits to await fetchData the closing paren and the semicolon carry no signal at
all. api surface is fixed by the library. keywords are fixed by the grammar.

now line that list up against a normal toolchain. prettier and black and gofmt destroy the
formatting entropy. linters enforce the naming conventions. minifiers rename the identifiers
outright. review comments change names. refactoring tools rewrite the constructs.

so in prose the watermarks adversary is a person deliberately paraphrasing to hide something. in
code the adversary is the build pipeline, its not adversarial at all, and it runs automatically on
every commit. gofmt isnt trying to strip anything. it just does.

which leads to the part i think matters more than the mechanism.

you make the two error directions clearly, watermark detected does not mean ai wrote everything
and no watermark does not mean a human did. thats right and most coverage skips it. but the two
are not symmetric in kind, theyre different types of statement. a detected watermark is positive
evidence that a specific model touched the text. an undetected watermark is not evidence of
anything, its just the absence of a reading. one is a signal, the other is a null result, and null
results get read as negatives constantly.

that asymmetry is survivable in prose. in code it inverts the instrument, because absence becomes
the default rather than the edge case. and the failures wouldnt even be randomly distributed.
theyd correlate with tooling discipline. a team with formatters and linters and a minified build
strips the signal as a side effect of being competent. a team committing unformatted code with
model chosen variable names keeps it. so a detector pointed at repositories would systematically
under report exactly where engineering practice is strongest and over report where it is weakest,
which is a bias with a direction, not noise.

not an argument against doing it. marking is better than not marking and the eu is right that
machine readability is the floor. but if the detector ever gets used for anything consequential,
licensing, provenance, academic integrity, the thing that has to travel with it is that it answers
one question and stays silent on the other, and staying silent is not the same as saying no.

id also like the docs. specifically whether they attempted code at all or whether the marking is
natural language only and claude code output is covered on paper rather than in practice.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Yes, exactly 😄 A new version of the docs is out now, and Anthropic basically admits that code is going to be a difficult case, especially for short snippets.

So we’ve reached a wonderfully useful state of EU compliance 😂

Watermark detected → AI touched this at some point, but we still don’t know how much of it was written by a human.

No watermark detected → we know… absolutely nothing.

And with code, where normal formatting, refactoring, linting, etc. can destroy the signal anyway, that second case is probably going to be especially common. So yes, I think your point about a null result being very different from negative evidence is exactly the problem here.

Collapse
 
mudassirworks profile image
Mudassir Khan

the statistical token angle is where practical questions pile up. if the signal requires paragraph length text to emerge, what happens when that paragraph gets edited down by 30% for a larger human doc — does the signal degrade proportionally or fall off a cliff? the editorial review exception in the EU Act seems to create a weird incentive structure: route AI output through nominal human editing, claim the exception.

the thing nobody's addressing is detection confidence thresholds. every real watermark system has false positives. what's the model for how platforms act on "probably watermarked" vs "confirmed"?

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

These are absolutely valid questions and unfortunately... we simply don't know yet! 😅 Anthropic says it will publish more technical documentation soon, so for now, all we can do is wait.

Interestingly, Anthropic itself explicitly mentions false positives and the possibility of the watermark becoming undetectable after editing or transforming the text, so they're definitely aware of these limitations.

I have a lot of questions about this too, especially around confidence thresholds and how much editing the signal can actually survive. Hopefully the technical docs will answer at least some of them!

Collapse
 
buildpilots profile image
king li

agree

Collapse
 
benjamin_nguyen_8ca6ff360 profile image
Benjamin Nguyen

Yep, I like the AI slop feature on LinkedIn

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Hahahaha exactly, very helpful feature 😁

Collapse
 
benjamin_nguyen_8ca6ff360 profile image
Benjamin Nguyen

yes, it is :)

Collapse
 
elsie-rainee profile image
Elsie Rainee

Great insights here 💡 Definitely bookmarking this for later! 🔖

Collapse
 
metaeth77 profile image
metaeth77

I understand professionalism, but there are many ways to do this.

Collapse
 
officialmailkr profile image
오피셜메일

워터마크와 AI 문체 감지를 같은 문제로 보지 않아야 한다는 구분이 유용했습니다. 특히 인간 검토와 편집 책임이 있는 경우까지 함께 짚어 주셔서 실제 콘텐츠 운영 기준을 세우는 데 도움이 됩니다.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

맞아요! 바로 그 점이 EU 규정에서도 중요한 부분입니다. AI로 텍스트를 생성하는 것 자체를 금지하는 것이 아니라, 사람이 내용을 검토하고 편집한 뒤 그 결과에 대해 편집 책임을 지도록 하는 것이죠.

다만 사람들이 실제로 이런 콘텐츠에 어떻게 반응할지는 또 다른 문제인 것 같아요. 😅

재미있는 점은, 이 댓글도 사실 100% 사람이 쓴 글이라는 겁니다. 제가 먼저 폴란드어로 직접 쓴 다음, AI를 사용해서 한국어로 번역했어요. :D

Collapse
 
voltradoc profile image
Dr Haina

That was interesting read 😀

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Thanks ☺️

Collapse
 
thegm26 profile image
George Michalakis

Are we aware on when more info are going to be officially shared with the public? Thanks again for this article!

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

It’s already out! 😄 And all the assumptions I made in this article turned out to be correct. For text, the watermarking solution will be SynthID. When it comes to code, though, it looks like it’s going to be much harder to solve.

Collapse
 
polterguy profile image
Thomas Hansen
Collapse
 
socar profile image
SOCAR

So all you have to do is replace couple of words, or change their or sentence order, in resulting text to get rid of the watermark... Kind of security by obscurity.

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Hahaha, exactly! And the longer I think about it, the more I suspect Anthropic's plan was basically: “Let's implement something so the EU gets off our back” 😅

The end result might be that the people who are most determined to hide AI usage will remove the watermark anyway, while ordinary users who aren't trying to hide anything will be the ones dealing with the consequences.

Collapse
 
aiscouter profile image
AI Scout

how refreshing finally good news :))

Collapse
 
daymondhyper profile image
DaymondHyper

Bookmarked. The part about evaluation being mandatory is exactly what I keep missing in my own experiments. How do you measure regressions, a separate test set?

Collapse
 
publiflow profile image
PubliFlow

Good coverage of ML patterns. I'd stress that monitoring data drift and model staleness is as important as the initial training — a model that was accurate at launch can silently degrade without proper observability.

Collapse
 
publiflow profile image
PubliFlow

Interesting ML/AI content. One practical consideration is the inference cost — model distillation and quantization can dramatically reduce deployment costs while maintaining acceptable accuracy for many use cases.

Collapse
 
royalx777 profile image
Royalx777 Apk

Explore my home post