DEV Community

Cover image for The "AI" Badge Doesn't Measure What You Think It Does
Pascal CESCATO
Pascal CESCATO Subscriber Community Curator

Posted on Edited on

The "AI" Badge Doesn't Measure What You Think It Does

Proves watermarks track provenance poorly

Anthropic signed the EU AI Act's Code of Practice on Transparency of AI-Generated Content, and started marking text produced by Claude with an invisible statistical watermark. Within days, the same event got read two completely different ways on my feed.

The first reading, on dev.to, is by @sylwia-lask

She actually went and read the official documentation and reported what Anthropic really says: the watermark doesn't prove a text was entirely AI-generated – you can write something yourself, have Claude correct it or translate it, and it can still carry the mark. The reverse holds too: no watermark detected doesn't prove a human wrote everything.

The second, on Medium, called the announcement a "nuclear bomb" for the AI-generation community, and claimed "you are effectively carrying a digital scarlet letter." No citation of the actual documentation. Just outrage, with the appropriate amount of dramatic punctuation.

Between the two, I left a comment under Sylwia's article. Written in French, translated by ChatGPT, signed off with a question: how do you classify a text thought through, written, and reviewed by a human, but whose final English phrasing came out of a model? It wasn't a rhetorical question. I still don't have an answer.

Three things people keep conflating

The Medium article mixes together, without ever saying so, three mechanisms that have almost nothing to do with each other.

Anthropic's watermark. A statistical signal baked into the model, documented, which openly acknowledges its own limits: it can disappear after enough editing or translation, and its presence says nothing about who came up with the idea in the first place.

Third-party consumer detectors. ZeroGPT and its relatives. Tools that have existed for years, unrelated to Anthropic, whose reliability has never been seriously demonstrated at scale.

Platform decisions. A badge displayed on Medium, a curation algorithm that penalizes flagged content. These are editorial choices specific to each platform, not a mechanical consequence of the watermark.

The Medium article writes as though the first link automatically triggers the next two: Anthropic marks, so detectors will catch everything, so platforms will punish. The reasoning collapses the moment you separate the links – and for good reason: detecting that a model was involved in producing a text is not the same as determining that the text was "written by AI." That conflation hides a deeper one, which public debate almost systematically ignores.

Assisted, generated, produced: three different verbs

A text assisted by AI stays under human editorial control from start to finish: the idea, the angle, the structure, and the final decision to publish belong to someone, no matter how many back-and-forths with a model happened along the way to correct, rephrase, translate, or pressure-test a line of reasoning. A text generated by AI comes out of a prompt, without that upstream control – the idea itself came from the human, but the generated text, as it stands, belongs to the model. A text produced by AI at industrial scale is something else again: an automated publishing pipeline, without anything resembling editorial oversight.

What separates the three, then, isn't how much the model intervened – it's who kept their hand on the decisions.

These three cases carry entirely different editorial responsibility. Treating them as a single category ("AI content") means judging a text on a binary criterion where the reality is a full spectrum of nuance – which is exactly what an undifferentiated "processed by AI" badge does.

What an actual test shows

I ran an article I wrote in April 2021 – before any consumer-facing LLM existed – through ZeroGPT. A test about a year ago gave it a 97% AI probability. A recent test, on the exact same text, with the exact same tool, gives 8.6%.

Same tool. Same text, down to the punctuation. Two incompatible scores, a year apart.

Digging into the flagged passages in the second test, a pattern emerges. These aren't random sentences: they're consistently the most neutral, most pedagogical, most well-structured passages in the piece – a definition of what a web server is, an explanation of CentOS Stream, a step-by-step automated update procedure. The passages where my voice actually comes through – the self-deprecation, the verbal tics, the Neapolitan moka pot bought at a flea market – are never flagged.

One test doesn't prove what a detection model measures in general. But this one strongly suggests the tool reacts to stylistic neutrality and structural regularity, not to a text's actual origin. Which is precisely the problem: those characteristics existed in human writing long before LLMs did. The detector is chasing a style whose origin is human, using the machine's imitation of it as the reference point. The reasoning eats its own tail.

How this piece was actually made

Since the whole point of this article is how hard it is to judge a text by its origin rather than its content, it's worth being transparent about how this one was produced.

The starting point wasn't this article: it was a comment under a Medium post that annoyed me enough to reply. From there, several hours of back-and-forth with Claude – not to have it write for me, but to pressure-test angles, check numbers, and go dig up original sources instead of relying on my own fuzzy memory. The ZeroGPT test wasn't improvised for the article, either: I reran it to verify what I was recalling from memory, and the result changed the conclusion I was about to draw.

First draft written in French – the language I actually think in. Several rounds of review, cross-checked across a few different AI models – ChatGPT, Kimi, Grok, Mistral – to catch inconsistencies and spots where the argument went soft. I then read and corrected it myself, outside of Claude (I use NotepadMD), because I never let a text out the door without going through it line by line. Translation into English after that, then a review of that translation, because a lexically faithful translation can still betray the tone if nobody checks it.

By the end of this process, the text probably carries a watermark. It also took up several hours of research, verification, and both automated and human review – and my own review isn't the least of it: after every rewrite, and again at the end of the process. Both facts are true at the same time, and no detector will ever be able to tell them apart.

What a badge actually costs

Everything above is a technical demonstration. It says nothing about what happens once the text is published, and that might be the more important part.

A "processed by AI" badge, applied without distinguishing assisted, generated, and produced, puts an author who spent hours thinking, checking, and writing in the same bucket as a content farm publishing a hundred articles a day with no human oversight. A rushed reader doesn't see the nuance – they see the badge, and file the thinking author under "cheater." That's a real loss for someone who never cheated: their work gets judged on a signal that measures neither the effort nor the actual editorial control behind it, only whether a tool showed up somewhere in the chain.

And that loss has an absurd mirror image. The same readers who reject an article marked "AI" accept, without blinking, the AI-generated summary sitting at the top of their Google search results – without reading it critically, without checking the sources it pulled from, and most often without ever clicking through to the original article. Independent studies converge on this: when an AI summary appears in a Google search, clicks to third-party sites drop by half, sometimes more, depending on methodology. In other words, the reader who calls an article "AI slop" over a badge probably let an AI summarize ten other topics for them that same week, without ever checking what it kept or what it distorted. AI isn't rejected on principle. It's rejected when it's visible and claimed, and accepted when it's invisible and imposed by default.

This shift from click to summary is also a broader loss of reading, independent of any badge: a full article, nuanced, sometimes contradictory, replaced by three smoothed-over lines nobody ever questions. The reader loses the friction that would have forced them to evaluate a source, a line of reasoning, a style – exactly what the badge claims to let people judge, and what the summary skips without any debate at all.

What this doesn't fix

None of this will stop an "AI slop" comment under this piece. I've never gotten one on dev.to. On Reddit, I have.

But look at who tends to write that kind of comment. Rarely people genuinely engaging with the topic. Often people chasing the buzz of a word that lands. On Reddit, the accounts behind these comments often show a broader pattern of reflexive, repeated downvoting across unrelated posts – a signal of rejection that has little to do with the content actually in front of them.

The technical demonstration and the social judgment are two different things, and the second doesn't get talked out of itself by the first. "AI slop" has stopped being a judgment about a text at all: it's become a blunt, reflexive rejection signal people display, regardless of whether they read anything.

If the badge doesn't measure origin, and social judgment doesn't care about the measurement, maybe the question worth asking isn't "who wrote this" but "who thought this through."

The nuance matters here. In the process I just described, the initiative never changed hands: I'm the one who opens the session, sets the topic, decides whether an angle holds up or needs to be dropped. A model doesn't come looking for me with an idea I didn't already have. And if it writes a paragraph I didn't ask for, it goes in the trash – which is perfectly observable in what I keep and what I discard from one session to the next.

Who thought through the initial disagreement with the Medium piece, and chose to separate watermark, detector, and platform instead of treating them as one thing? Who decided a test was worth more than an opinion, and reran it twice instead of republishing a fuzzy memory?

Those are questions I can answer, text by text, decision by decision. No detector asks them. It looks at the final shape and infers an origin from it – when the origin was never in the shape. It was in the chain of decisions that came before it.

Top comments (124)

Collapse
 
sylwia-lask profile image
Sylwia Laskowska

Great read, Pascal, as usual! I just want to add that apparently open-source just "fixed" the watermark problem 😅 They're blazing fast 🤣 github.com/guillaumemeyer/watermar...

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Awesome! Your article inspired me, so this one came naturally… 😄

Collapse
 
buildbasekit profile image
buildbasekit

The part that stands out to me is that the final output tells us very little about how it was created.

A developer can use AI for 20% of the work or 80%, but still be the person making the important decisions.

For me, the bigger question is not “Was AI involved?”

It’s “Could you explain and defend what you shipped without it?” 😅

Collapse
 
weirdcodesofficial profile image
Weird Codes

Shoudn't i make my devlog using ai, if i can explain what i build?

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Absolutely. 😄

If you built the thing, understand the decisions behind it, and can explain and defend them, I don't see why using AI to help turn your notes into a readable devlog would make the devlog somehow less yours.

In that case, the AI is helping with the expression of your experience, not replacing the experience itself.

That's precisely why “AI involved: yes/no” is such a poor proxy for authorship.

Thread Thread
 
weirdcodesofficial profile image
Weird Codes • Edited

I used AI for devlog coz i used AI for writting the code for my game. Other than the code i decided everything what to keep what to reject from the code and the other parts of the game. Hence i felt kinda lazy to write devlogs too coz i found the same AI wrote a good devlog as it wrote the code itself. I am not a traditional dev, i thought AI must be used if not want to waste time as other professional devs also use AI now days.

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

That actually makes your case even more interesting.

If AI generated the code and the devlog, but you decided what the game should be, evaluated the generated code, rejected or kept parts of it, and made the other design decisions, then the interesting question isn't whether AI was involved. It obviously was.

The question is whether the devlog accurately represents your decisions and your development process.

And I wouldn't call that necessarily lazy. If writing the devlog isn't the part of the process you want to spend your time on, using a tool to document it can be perfectly rational.

The thing I'd be careful about is making sure the AI-generated devlog doesn't invent reasoning you never had just because it makes a better story. 😄

That's actually another provenance problem: the difference between documenting what you decided and generating a plausible story about why you decided it.

Thread Thread
 
weirdcodesofficial profile image
Weird Codes • Edited

I didn't use AI generated devlog for plausible story for sure coz i don't remember if i missed any devlog posted without my own confirmation, i deleted texts and unwanted things many times from devlogs before posting, but i still liked the way how quick AI generated the entire devlog for me that could take my expensive time if i would do that entirely by me. And yes, i found AI much better for formating the text in markdown much better than myself.

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

That's exactly the kind of workflow I would call AI-assisted rather than AI-authored.

The AI did the expensive part in terms of time: turning your development history into a coherent Markdown document and handling the formatting. But you still acted as the editor — deciding what stayed, what was removed, what was corrected, and what was actually published.

And honestly, Markdown formatting is a perfect example of why I don't think “AI was involved” tells us very much. If a tool saves you an hour of formatting work, that doesn't suddenly make the underlying experience or decisions belong to the tool.

The important part is that you remained the gatekeeper of the final artifact. That's a much more meaningful distinction than whether AI touched the text. 🙂

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Exactly. And I think that's a much more useful question than “Was AI involved?”

I might even push it one step further: Can you explain, challenge, modify, and defend what you shipped — without outsourcing the responsibility for it to the tool?

That doesn't prove who wrote every line or sentence. But it tells us something much more relevant: whether the human who put their name on the result actually owns the decisions behind it.

A watermark can tell you that a tool participated. It can't tell you whether you understood what you shipped. 😅

Collapse
 
earlgreyhot1701d profile image
Earl Grey

"Assisted, generated, produced" - Great breakdown @pascal_cescato_692b7a8a20 AI helps us code, helps us write, helps us find shortcuts. The thoughts that go into it, our style of writing, our voice, what problem we choose to explore, well those are uniquely human and are ours to keep, not to be counted. Watermark or not.

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO • Edited

Exactly. The tool can participate in the execution without becoming the origin of the thought.

That's why I find “assisted / generated / produced” more useful than a binary “AI / human” label. The important question isn't whether a tool was involved, but where the human decisions, intent, and responsibility remained.

Watermarking can label an artifact. It can't measure that.

It's a bit like “trafilato al bronzo” on a pasta package: true, measurable, and completely insufficient to tell you whether the wheat was good or whether the pasta was properly dried. 😄

Collapse
 
klaudiagrz profile image
Klaudia Grzondziel

I also like this distinction! 💯 As a technical writer who reads a lot of text every day, I have nothing against AI-assisted content, and I think this is just how people work nowadays. Unfortunately, I've seen enough AI-generated or AI-produced docs with no human oversight to know how harmful they are. This is what I call slop.

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Exactly. That's probably the distinction I care about most.

AI-assisted content with a human actually reviewing, correcting, restructuring and taking responsibility for the result is one thing. AI-produced content published with little or no human oversight is another.

The problem is that a binary “AI” label collapses both into the same category, while the quality difference can be enormous.

So for me, “slop” isn't a synonym for “AI-generated”. It's what happens when production is automated but editorial responsibility disappears.

Thread Thread
 
edmundsparrow profile image
Ekong Ikpe

For me, it's a fugazi. Believe the watermark at your own discretion.

Take the case of PDF:
provenance metadata can be lost or altered across format transformations.

Until a PDF can be reversed back to its original provenance, forgerrit.

Watermark ≈ certificate of participation
Exam/project defense ≈ evidence of competence

Are we testing knowledge, or are we testing the provenance of the learner's cognitive process? 🤔

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

The real question is whether we're evaluating the artifact, the competence behind it, or the provenance of the process that produced it. Those are three different things. A single stamp can't reliably answer all three.

Collapse
 
mudassirworks profile image
Mudassir Khan

the 'three mechanisms people keep conflating' breakdown is doing the real work here. the provenance chain (Anthropic watermark → ZeroGPT → platform ban) assumes each link automatically fires the next, and the assumption breaks at every joint.

the harder question your French comment gestures at is edit distance. if i draft every paragraph, run Claude over it for sentence rhythm, and review the output: who wrote it? the watermark says Claude was involved. the reader got my thinking. the platform algorithm has no field for that.

the assisted category is the one nobody has a defensible definition for yet. have you found any platform that draws the line somewhere that actually holds?

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

That's exactly the problem I keep running into: I haven't found a boundary that remains defensible once you move away from the extremes.

“AI wrote the whole thing” is relatively easy to classify. So is “I used AI to translate this.” But between those two, you get an almost continuous spectrum: brainstorming, outlining, restructuring, rewriting, sentence-level editing, tone adjustment, translation, fact-checking, etc.

And edit distance doesn't solve it either. I could write 100% of the ideas and drafts, have Claude rewrite every sentence, and still be the person who decided what the piece says, what evidence matters, and what gets published.

That's why I think “assisted” is such an uncomfortable category: there's no obvious threshold where assistance suddenly becomes authorship.

I haven't yet seen a platform define that boundary in a way that survives those edge cases. Most seem to need a binary field because binary fields are easy to implement — not because the underlying reality is binary.

Collapse
 
mudassirworks profile image
Mudassir Khan

the 'binary fields are easy to implement' observation is the actual story. the platform draws the line because their database doesn't support a slider — not because they believe in the line.

the one signal that might hold up: revision history. if you wrote, deleted, and rewrote over 40 mins, that behavioral trace is harder to fake than a checkbox.

any chance a platform like substack or medium has started collecting that signal even if they're not surfacing it yet?

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

That's actually a much more interesting signal than a binary checkbox. Revision history gives you a trace of the process rather than an inference from the final artifact.

Medium already has revision history, and Substack keeps draft versions as well, so the raw material exists. But I haven't found evidence that either platform currently uses that history as a provenance signal for AI involvement.

And even revision history wouldn't prove “human authorship” by itself — you could paste an AI-generated draft as your first version. But it could provide much richer evidence about the actual process than a detector score ever could.

Thread Thread
 
mudassirworks profile image
Mudassir Khan

yeah the paste problem is the real edge case. but platforms could track it. a 2000 word first save with zero preceding keystrokes is still a behavioral fingerprint even if it looks like a draft.

the richer question: would platforms ever surface that data to writers about their own work, let alone use it as a moderation signal?

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO • Edited

Exactly. And there's another problem: I can write my draft in my own editor, with my own tools, and paste the whole thing into the platform.

If a platform starts treating “2,000 words pasted with no preceding keystrokes” as a provenance signal, we're getting dangerously close to monitoring how people work rather than evaluating what they produced.

At that point, we're not far from Orwell's telescreen: the platform isn't just judging the artifact anymore — it's watching the writer's behavior to decide whether the artifact is legitimate. 😄

I'd rather have transparent provenance that the author controls than an invisible behavioral score they never get to see.

Thread Thread
 
mudassirworks profile image
Mudassir Khan

fair — and the telescreen comparison holds once keystroke cadence becomes the signal. the cut i keep returning to though: revision history or published diffs are behavioral signals the author can see and contest. keystroke cadence scraped server side is behavioral data they'll never know was collected.

same underlying action (writing behavior), completely different accountability surface. do you think platforms would ever ship the first version openly vs quietly roll out the second?

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO • Edited

I think the first one is much more defensible precisely because the author can see it, correct it, and contest the interpretation. It becomes part of the provenance they control rather than a hidden score used against them.

The second one is where I'd draw the line. Once a platform silently collects behavioral telemetry and turns it into a reputation or moderation signal, the writer isn't just being evaluated anymore — they're being profiled.

And that's the irony: a system supposedly designed to restore trust could undermine it by making the trust mechanism itself opaque. If platforms ever go down that road, I hope they'll at least publish what signals they collect and how they're used.

At the end, a better signal doesn't necessarily make a better system if the person being judged cannot see it, understand it, or contest it.

Thread Thread
 
mudassirworks profile image
Mudassir Khan

the 'contestable by the person being judged' framing maps onto GDPR's right to explanation — but for editorial reputation rather than credit decisions.

the part that would change the calculus: if platforms go opaque and a high profile writer gets false flagged without recourse, that's the litigation trigger. the pressure won't come from devs complaining.

do you think that's the only lever that would actually force platforms to publish what they collect, or are there softer paths first?

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

The softer lever may be transparency pressure: disclosure, auditability, and the ability for users to see and contest the signals attached to their work. But that only works while platforms have an incentive to cooperate. Once an opaque signal starts affecting visibility or reputation and a high-profile writer is falsely flagged, the calculus changes. At that point, litigation or regulatory pressure becomes much harder to ignore.

And ironically, the strongest trigger may not be AI detection itself, but undisclosed behavioral data being used to make reputational decisions. That creates a much cleaner accountability problem than “your AI detector got it wrong.”

Thread Thread
 
mudassirworks profile image
Mudassir Khan

the 'behavioral data used for reputation decisions' framing is the sharper legal theory. GDPR already covers this — Article 22 kicks in when automated decisions have 'significant effects,' and reputational scoring qualifies.

'AI detector wrong' is hard to litigate because there's no agreed ground truth. 'you used undisclosed engagement signals to deprioritize my content' is a data processing claim. much more concrete.

do disclosure mandates get there before litigation, or does the first high profile case need to happen first?

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

I suspect disclosure mandates can get there first, but probably not because platforms suddenly become transparent. More likely because regulators force them to explain what categories of behavioral data they collect and how those signals are used.

The interesting threshold is when “engagement optimization” becomes a decision with a measurable effect on someone's reach or reputation. At that point, the distinction between a ranking signal and a reputational score gets very thin.

A high-profile case could accelerate that process dramatically, though. Once there is a concrete plaintiff and a demonstrable undisclosed signal, the argument stops being theoretical: what did you collect, what decision did it influence, and where did I consent to that processing?

That seems like a much harder question for a platform to answer than “prove your AI detector is accurate.”

Collapse
 
alifunk profile image
Ali-Funk

Well done all around
You did a great job putting this article together.
The structure, the angle,the thought process behind it.
I am happy I came across this article of yours.

This wasn't AI generated. Real human watermark 😃

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Thank you! 😄 And now I have to ask: where do I get my “real human watermark” verified? 😂

Collapse
 
alifunk profile image
Ali-Funk

Hard to say...I would go for a real human blood test and have the family tree checked. #Ex Machina
Can´t be to carfeful these days, am I right ? ^^

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

😂 Careful — that's how it starts. First they ask for a human watermark, then a blood test, and before you know it you're sitting in a glass room being tested for consciousness.

I think I'll stick with “I can explain and defend what I wrote.” 😄

Collapse
 
gnomeman4201 profile image
GnomeMan4201

I think there’s a flip side to the false-positive problem Dean raised: once the badge has consequences, it doesn’t just risk flagging the wrong people it changes who is willing to leave the signal intact.

A careful writer using AI for translation, editing, or research may have no reason to conceal that provenance. Someone mass generating content, though, has every incentive to learn what triggers the badge and route around it.

So over time the signal could become strangely inverted: the people acting transparently remain easiest to identify, while the behavior the system is actually trying to discourage becomes increasingly optimized to evade detection.

That starts looking like a Goodhart’s law problem. Once the proxy affects ranking, monetization, or reputation, people optimize against the proxy rather than the underlying behavior.

It makes me think the design question isn’t only “how accurate is the badge?” but “how do you make honest disclosure less costly than concealment?” Otherwise the system may select against the transparency it was supposed to create.

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Yes — and that's the part I find most dangerous. Once the signal has a cost, it becomes a target for optimization. The transparent user keeps the signal; the user who has the strongest incentive to hide AI involvement starts optimizing around it. So the system can gradually select against transparency rather than encouraging it.

And that's a much deeper problem than false positives: the measurement itself changes the population being measured.

A badge designed to reward transparency can end up making transparency the most expensive option.

Collapse
 
gnomeman4201 profile image
GnomeMan4201

One thing this makes me wonder about is selection bias over time. Once transparent users are disproportionately represented among detectable cases, any statistics derived from badge prevalence become increasingly misleading. The system wouldn’t merely punish the wrong population; it could start producing data that appears to justify its own conclusions because the evasive population is systematically missing.

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Exactly. And at that point the badge isn't just a noisy measurement — it's sampling a population that has already been shaped by the measurement itself. The detectable cases become a biased sample, while the successful evasions disappear from the dataset.

That creates a nasty feedback loop: the more the badge is used, the less representative its observed population may become, while the resulting statistics can make the system look more accurate than it actually is.

Goodhart's law was already bad enough. Add survivorship bias and you get a very convincing illusion of evidence. 😄

Collapse
 
heinrichneb profile image
Heinrich Neb

The thread has already converged on something I'd like to add two measurements to — because we've been running exactly what several people here are describing, and both numbers are unflattering.

For about three months we've kept a decision record alongside the work: what was tried, what failed, which files it touched, which earlier decision it contradicts. Captured while the decisions happen, not reconstructed from the artifact. Two things we didn't expect:

  1. The human field rots first. Of 493 recorded decisions, 298 carry an author and 195 do not. The evidence, the commands, the files — all reliably captured, because the tooling writes them automatically. "Who decided this" was the one optional field, so two fifths of it is simply gone. If provenance is going to be part of the architecture rather than an audit trail bolted on, the human-decision field cannot be the one that's easiest to skip. Ours was, and it's the only field we actually care about.

  2. Capturing it isn't enough — the retrieval shape decides whether it changes anything. We had a decision whose critical fact sat at character 323 of the record. The summary view showed the first 100 characters. The record was on screen, correctly captured, correctly linked to the right file — and the same mistake happened anyway. Provenance that is stored but not surfaced at the moment of the next decision behaves exactly like provenance that was never captured. That one cost us a day, and no amount of better capture would have prevented it.

So a question rather than a claim, for you and for @suraj09 with the living-graph idea: is anyone measuring not just what gets recorded, but whether a recorded decision actually changed a later one? That seems like the only number that would prove the whole approach — and it's the one I can't get at yet.

Collapse
 
suraj09 profile image
Suraj Suradkar

That’s a really useful question, and I think it exposes the harder part of the living-graph idea.

I don't think capture volume is the right success metric either. I’d want to measure something closer to decision influence: was an earlier decision/evidence actually surfaced at the point of a later decision, and did it change the action, constraint, or outcome?

The tricky part is attribution. A later decision can change for many reasons, so simply linking two records isn't enough. You’d probably need to distinguish retrieved → considered → influenced → changed rather than treating retrieval itself as proof of impact.

And your 323-character example is especially important. If the relevant part of provenance isn't surfaced when the next decision happens, then operationally it behaves like missing provenance.

I think that “did this knowledge actually change what happened next?” metric is much closer to the thing worth measuring.

Collapse
 
heinrichneb profile image
Heinrich Neb

Your ladder is doing something mine wasn't, and I want to name it before adding anything: I had been treating "retrieved" as the floor and quietly letting it stand in for "arrived". Splitting it into retrieved → considered → influenced → changed is exactly where my Tier 1 was doing work it hadn't earned.

One genuine question about the middle two, because I've failed at it and would rather be wrong: is there a way to observe "considered" and "influenced" that isn't the attribution problem in a new coat? Every approach I've tried ends the same way — I find a plausible story linking record and outcome, and then I have no test for the story. If someone has a real handle on those two rungs, I'd rather learn it than route around it.

Because the routing-around I've been trying is a bit of a dodge, and I'm curious whether it survives contact with you two. Instead of asking "did this record change this decision" — which needs a counterfactual for a single event — ask: over N occasions of the same situation, did the failure recur at a different rate? Nothing gets attributed to any single decision, so there's nothing left to attribute. What it costs is the ability to point at one case and say "that one", which happens to be the claim I couldn't defend anyway. Is that a real escape, or have I just moved the problem somewhere it's harder to see?

And now the part I'm oddly pleased about. I have to correct myself, and the correction turned out more useful than the original claim.

I said there were three truncation points between our store and the model. I went looking for a fourth on that path and didn't find one. I found it somewhere else entirely — on the path out, to the human. Our "export everything" command was being served by a dashboard summary endpoint: 493 records in the store, the export wrote 50, each cut at 120 characters, mid-word.

So it isn't a fourth truncation point on the model path, and I don't want to dress it up as one. It's the same shape on a path I hadn't thought to check — which is worse, and more interesting. My enumeration was scoped to a route rather than to the pattern. And every layer was individually correct: the summary endpoint was built for an IDE list and is right for that, the export command came later and reused what was already there. Nobody made a mistake. Two questions just quietly shared one answer.

Hence the thing I'd actually offer, which is the shape rather than the bug: these appear at every boundary where one component asks another for "the data" and gets back a view built for somebody else's question. Enumerating three of them and testing exactly those three is how I ended up confidently wrong in public earlier this week.

That also makes the Tier 2 test cheaper and more portable than I'd described it, so take this if it's any use:

# Tier-2 canary: does the decisive fact survive the trip to the model?
# Assert on the bytes you actually SEND — not on what retrieval returned.

canary = "ZX9-" + uuid4().hex[:8]
store_record(body=filler(300) + canary + filler(300))   # plant it DEEP

sent = []
client = httpx.Client(event_hooks={"request": [lambda r: sent.append(r.read())]})
# the openai and anthropic python SDKs both accept http_client=client

run_the_thing_that_should_retrieve_it()

assert any(canary in b for b in sent), \
    "retrieved fine — truncated somewhere between the store and the wire"
Enter fullscreen mode Exit fullscreen mode

Two details decide whether this tests anything. Assert on the outbound request body, not on your retrieval layer's response — ours returned HTTP 200 with the correct record attached at every single stage while the model was shown a preview. And plant the canary deep: at character 20 it passes everywhere and proves nothing. The useful property is that it doesn't care how many boundaries sit in between, or whether you knew they were there — which, going by this week, you don't.

Last thing: I asked a question upthread and then didn't answer it myself, which isn't fair. Yes, I'd accept a hold-out as evidence, and I don't think I'd accept anything weaker. What I'm genuinely unsure about is whether I'd accept one from a system I hadn't built, without seeing how the eligible turns were chosen — I suspect that's where the result gets decided, long before any numbers come out. How would you want that part shown?

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

This has gone considerably further than I expected when I wrote the article. 😄 I'm going to let you two take the provenance rabbit hole from here — but the distinction between “retrieved” and “actually received by the model” is a very good example of why I wrote the article in the first place: the observable signal is often not the thing we think we're measuring.

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

I think we're now exploring a question beyond what I was trying to address in the article, but that's exactly what makes this thread interesting. 😄

My original point was simply that provenance cannot reliably be inferred from the final artifact. What you're describing is the other side of that: once provenance is captured, you still have to prove that it remains usable and actually influences subsequent decisions.

Those are two different problems — and probably worth keeping separate rather than creating another proxy that we mistake for the thing itself.

Otherwise we may end up building a very sophisticated measurement system for measuring the wrong thing.

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Those are excellent observations — especially because they expose two different failure modes. Capture can fail at the point of authorship, and retrieval can fail even when capture is perfect.

And I think you're right about the metric. "How much provenance did we record?" is almost meaningless by itself. The interesting measurement is whether a recorded decision changed a later decision, prevented a repeated mistake, or caused someone to revisit an earlier conclusion.

That's also where the living-graph idea becomes more than a storage model: the value isn't in preserving the graph, but in making the relevant part of it intervene in the next decision.

I don't have a clean metric for that yet either. But "decision changed because of prior provenance" sounds much closer to the outcome we should actually be measuring.

Collapse
 
heinrichneb profile image
Heinrich Neb

That framing — "decision changed because of prior provenance" — is the one I'd want to build against too. Let me try to make it operational, because I think it splits into three tiers, and only the third actually answers the question.

Tier 1 — was it surfaced? Did the retrieval return anything at all for this moment? Cheap, and we measure it: 64 % of our recalls come back with something. Nearly worthless on its own, but it's the floor.

Tier 2 — did it survive the trip? This is the one I underestimated, and it's where the character-323 story actually lives. There are three truncation points between our store and the model: a 500-token cap on the record, a summary window, and a 100-character line in the briefing. Each is individually defensible. Nobody had ever tested end to end that the specific fact survives all three. A pipeline can pass every unit test — right record found, citation attached, HTTP 200 — while the model is quietly shown a truncated preview. So Tier 2 is: does the decisive sentence appear in the bytes that actually reach the model? Mechanical, testable, and we're building that test this week.

Tier 3 — did it change the outcome? I don't think this is measurable by observation, only by withholding. Suppress injection for a small random share of eligible turns, then compare how often the same failure recurs between the two groups. It's the METR design applied to provenance, and it has an uncomfortable property: you have to deliberately make the system worse for part of your sample. I suspect that's the real reason nobody has this number — not that it's hard to compute, but that nobody wants to pay for the control group.

One thing that makes Tier 3 cheaper than it sounds: "the same mistake recurred" doesn't need human judgement if you already detect contradictions or duplicates. A second record written for a topic that already had one is the repeat, mechanically.

So my honest position: Tier 1 is vanity, Tier 2 is the engineering work almost everyone is skipping, and Tier 3 is the only proof — and it costs a control group. Curious whether you'd accept a hold-out as evidence, or whether you'd want something that doesn't require degrading part of the system.

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

I think we've reached an interesting point where the discussion is starting to move beyond the scope of my article. 😄

The connection I was making was much narrower: a watermark or detector looks at the final artifact and tries to infer something about its origin. Your work is about capturing the decision process itself, which is almost the inverse approach.

The Tier 1/2/3 framework is interesting, but it raises a different question: whether a provenance system is operationally effective once you have decided to build one. My article is mostly questioning whether the artifact alone can provide the kind of provenance signal we're asking it to provide in the first place.

And I think that's an important distinction to keep, otherwise we end up measuring the wrong thing again. 🙂

Collapse
 
suraj09 profile image
Suraj Suradkar

@heinrichneb I think your N-occasion approach is actually more defensible. Instead of claiming “this record caused this decision,” we can ask whether access to accumulated knowledge measurably reduces recurring failures across comparable situations.

I’d treat retrieved → considered → influenced as observability signals, not causal proof. The hold-out comparison is where the stronger evidence comes from.

And your boundary-canary example is a great point: validating what retrieval returned isn't enough. We should test what actually crossed the boundary to the model. That’s a much stronger definition of “the model received the evidence.”

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Yes — and I think that's the useful distinction. Observability can tell us what happened in the pipeline; it doesn't automatically tell us what caused the outcome.

That's also why I like the boundary-canary idea: before asking whether evidence influenced the model, we should at least be able to prove that the evidence actually reached it. Otherwise we're trying to measure an effect of something the system may never have seen.

And this brings us surprisingly close to the original point of my article: in both cases, the dangerous step is treating an observable proxy as if it were the thing we actually care about.

Collapse
 
heinrichneb profile image
Heinrich Neb

"Observability signals, not causal proof" is the phrasing I was missing — I'll
steal it, with attribution. It also settles my discomfort with the middle
rungs: "considered" and "influenced" don't need to be proven, they need to be
logged, and the proof lives one level up, in the hold-out.

So let me make that concrete, because I'd rather be held to something: we're
going to run it. The part I'd value your eyes on is the eligibility rule. My
current draft: a turn is eligible if retrieval returned a record above
threshold AND the canary confirms it reached the model; a recurrence is
mechanical — a second record written for a topic that already had one. Both
rules published before the run, alongside the split ratio. Is a pre-registered
rule like that enough for you to trust the resulting number, or would you want
the raw turn list as well?

And @pascal_cescato_692b7a8a20 — you've now twice watched us dig a provenance
tunnel under your watermark article, which is a hospitality I don't want to
overstretch. 😄 When the hold-out numbers exist, I'll write them up as their
own piece instead of as comment #47 here. Thank you for hosting the start of
it — "an observable proxy mistaken for the thing itself" turned out to be the
sentence both problems share.

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

😂 At this point, I think the provenance tunnel has officially become its own project. And that's probably a good thing.

I'm glad the “observable proxy mistaken for the thing itself” idea proved useful beyond the original watermark discussion. That's exactly the kind of rabbit hole I hoped the article would trigger.

For the experiment, I wouldn't pretend to be the right person to validate the methodology — but pre-registering the eligibility and recurrence rules before seeing the results sounds like a very good way to avoid moving the goalposts afterwards.

And yes: please write the results as their own piece. I suspect comment #47 would be a terrible place to hide them. 😄

Collapse
 
suraj09 profile image
Suraj Suradkar

@heinrichneb I’d want the raw turn list as well, but mainly for auditability rather than as the primary result. The pre-registered eligibility and recurrence rules give you the strongest protection against moving the goalposts.

Ideally, publish the rules + split ratio first, then the eligible-turn IDs/list and exclusions after the run. That makes it possible to verify that the population actually matched the rule without exposing unnecessary content.

The key for me would be: can someone reproduce the cohort selection from the published rule and raw identifiers without knowing the outcome first? If yes, I’d consider that a pretty strong setup.

Collapse
 
suraj09 profile image
Suraj Suradkar

This is a really interesting distinction: the origin of a piece of content isn't necessarily visible in the final output.

It makes me think the more useful signal isn't “was AI involved?” but “what was the chain of decisions behind this?”

Who chose the direction, what evidence was checked, what was rejected, and what ultimately made it into the final version.

In a way, that feels like a provenance problem rather than a detection problem. The final artifact alone doesn't contain enough information to explain why it should be trusted.

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Exactly. That's the distinction I was trying to get at.

If provenance matters, it has to be captured during the process: what was considered, what was rejected, what was verified, and which decisions were ultimately made by the human.

Trying to reconstruct all of that from the final artifact is fundamentally different from recording it as it happens.

Detection asks “what does this output look like?” Provenance asks “how did we get here?” Those are not the same question.

Collapse
 
suraj09 profile image
Suraj Suradkar

Exactly. I think that also changes how we should architect AI-assisted systems.

Provenance shouldn't be something we try to reconstruct from the final artifact later. The important events need to be captured as the work happens: what was considered, what was rejected, what was verified, and why the final decision was made.

Otherwise, months later, we may have the final code but lose the reasoning that makes that code trustworthy.

That’s one of the areas I’m exploring with Xeyria: treating decisions, evidence, and their relationships as durable project knowledge, rather than just storing the final output.

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

That distinction between storing outputs and storing the reasoning around them is exactly where I think this gets interesting.

A final artifact is the conclusion of a process, but without the rejected alternatives, evidence, decisions and context, you lose much of the information needed to understand why that conclusion should be trusted.

And I like the idea of treating those relationships as durable project knowledge rather than as an audit trail bolted onto the end. That makes provenance part of the architecture instead of a label added after the fact.

I'm curious to see where you take Xeyria with this. It feels like a much more meaningful direction than trying to make the final artifact confess where it came from. 🙂

Thread Thread
 
suraj09 profile image
Suraj Suradkar

Thanks Pascal. I think the interesting challenge now is making those relationships useful, not just preserving them.

For example, if a decision was based on evidence that later changes, the system should be able to surface that connection instead of treating the decision as permanently correct.

So I’m increasingly thinking of provenance as something closer to a living graph: decision → evidence → alternatives → outcome, with enough context to understand when an old conclusion should be questioned.

That’s the direction I’m exploring with Xeyria. 🙂

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Yes — and that makes it much more interesting than provenance as an audit trail.

The moment evidence can change, provenance becomes temporal: a decision isn't simply “correct” or “incorrect”; it was justified by a particular state of knowledge at a particular point in time.

A living graph could therefore answer not only “why was this decision made?” but also “what has changed since it was made that might invalidate it?”

At that point, provenance stops being a label attached to an artifact and becomes part of the knowledge system around it.

That's a very different — and much more useful — way of thinking about trust. 🙂

Collapse
 
glenallen profile image
Glen Allen

The distinction between AI involvement and AI authorship is probably the most important point here. A provenance signal can tell us that a model participated somewhere in the process, but it can't tell us who made the meaningful decisions. I think evaluating the chain of human decisions and accountability would be far more useful than treating AI involvement as a binary label.

Collapse
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Exactly. We have a remarkably precise label for the least interesting question: “Was a model involved?” And almost no signal for the questions that actually matter: “Who decided? Who verified it? Who is accountable?”

Collapse
 
glenallen profile image
Glen Allen

Exactly. I think the more useful signal would be a provenance chain rather than a binary badge what tool was involved, where human decisions were made, what was independently verified, and who ultimately approved the result. That wouldn't tell us everything about authorship, but it would give readers something much closer to the accountability they actually care about.

Thread Thread
 
pascal_cescato_692b7a8a20 profile image
Pascal CESCATO

Yes — and I think the important shift is from “AI provenance” to “accountability provenance.”

Knowing that a model participated is useful context, but knowing what it did, what the human changed or rejected, what was independently verified, and who ultimately stood behind the result is much more informative.

It still wouldn't give us a perfect definition of authorship — but perhaps that's the point. We don't necessarily need to prove who “wrote” every sentence. We need enough provenance to know who thought, judged, verified, and ultimately took responsibility for it.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.