DEV Community

Cover image for I gave an agent my posting history. It found a promise I never made.

I gave an agent my posting history. It found a promise I never made.

Eugeniya Ivanova on September 08, 2026

Every social platform rewards the same thing, and it isn't a secret: showing up regularly. The hard part isn't writing. It's deciding what to say t...
Collapse
 
mk023 profile image
Marco •

Really interesting failure sequence, Eugeniya. I think all three cases share the same deeper problem: the evidence was real, but the meaning assigned to it was stronger than the evidence justified. 🔍

The Zapier posts are a great example. The agent correctly observed that those records existed, but “you published this to your audience” and “you promised a follow-up” were additional claims that required provenance the post history did not contain.

So I really like the rule:

observation != interpretation != fact

I would almost want every historical item the agent consumes to carry context such as source, audience, purpose and confidence before it is allowed to reason from it.

The cross-platform case is the same problem from the opposite direction. The agent did not hallucinate anything, it simply had an incomplete evidence boundary and no way to know that the missing context existed elsewhere.

And the slop checker is probably my favorite failure here. It passed because it was measuring properties inside the draft, while the thing your colleague noticed depended on context outside the draft: reuse, recent history and the cumulative effect of individually reasonable patterns.

So the checker was not necessarily wrong. The property being asked of it was larger than its observation surface.

That makes the final “stop at the draft” boundary especially important. The model can propose, the checker can catch a subset of problems, but neither gets publishing authority simply because both are green.

I think the architecture starts looking like:

history -> provenance/context -> inference -> human confirmation -> draft -> local checks -> publish decision

Really good example of why agent memory becomes much more useful when it remembers where a fact came from and what it actually proves, not just the fact itself. 🧠🔐

Collapse
 
eugeniya_ivanova_4a58eadc profile image
Eugeniya Ivanova •

"The meaning assigned to it was stronger than the evidence justified" — that's the sentence I wish I'd opened the article with. All three were three stories to me; you've written them as one, and you're right.

The provenance idea is where it gets practically interesting. Attaching source, audience, purpose, confidence to each historical item is exactly what would have stopped the Zapier misread at the root — a QA run tagged audience: none, purpose: test never gets read as a promise. The catch is that provenance has to be captured at write time, by the thing that created the record, and most systems don't. My test posts went into the feed with no marker because nothing along the way had a reason to add one. So the architecture is right, but it pushes the cost upstream: the feed has to start carrying context it currently throws away, and retrofitting that onto records that already exist is the hard part.

Your read on the slop checker is more generous than mine and also more correct. "The property being asked of it was larger than its observation surface" is the honest version of what I called a failure. It answered the question it could see. I was holding it responsible for a question that lived outside the file, which isn't the checker's fault, it's a category error on my part about what that tool is for.

And the pipeline you drew is the one I backed into without naming: history → provenance → inference → human confirmation → draft → local checks → publish. The one thing I'd underline is that the human-confirmation step and the publish-decision step are different gates doing different jobs — one asks "is the premise real," the other asks "should this go out," and collapsing them is how "both checks are green" quietly becomes publishing authority. Neither green light is the same as a decision.

Collapse
 
liammartin profile image
Liam Martin •

This is a fascinating breakdown, especially the part about structural "slop". You're totally right that vocabulary checkers are fighting yesterday's war. The real AI tells are the predictable rhythms: the dramatic single-sentence paragraph, the neat paradoxes, and those perfectly balanced three-item lists.

Since you built the logic with Claude, I'm curious about how you are handling the prompt for the actual draft phase now. Have you tried passing negative constraints to specifically ban those structural clichés (e.g., "Do not use three-item lists" or "Avoid dramatic one-line paragraphs"), or does Claude end up ignoring them when trying to mimic your voice? Would love to know if negative prompting worked for your use case!

Collapse
 
eugeniya_ivanova_4a58eadc profile image
Eugeniya Ivanova •

Great question, and the honest answer is that negative constraints in the draft prompt didn't carry the load for me, so I stopped leaning on them. "Don't use three-item lists" works right up until the model is also trying to match my voice and hit a length — then the structural habit slips back in, because it's not reaching for a list on purpose, it's falling into a rhythm, and a rule it isn't actively thinking about doesn't fire.

So I moved the anti-slop work out of the draft prompt and into a separate pass that runs after. Generate first, then check against the tells as its own step, rather than asking the model to avoid them while it's busy doing three other things. The catch — which is basically the whole point of that article — is that a checker only sees what's inside the text, so the structural pass still misses the cross-post repetition and the "four defensible choices add up to a machine" problem. Those I still catch by reading.

Short version: negative prompting helped a little, a post-draft check helped more, and neither one replaces a human read. Have you had better luck getting the constraints to actually stick in-prompt?

Collapse
 
liammartin profile image
Liam Martin •

That makes perfect sense. I've had similar struggles where negative constraints work for a sentence or two, and then the model completely forgets them by paragraph two! Treating the drafting and the editing as two distinct steps instead of one mega-prompt is a brilliant workaround. Thanks for sharing the behind-the-scenes on this, it’s definitely given me some ideas for my own workflows!

Collapse
 
prpatel05 profile image
Pratik Patel •

The Zapier validation posts are a clean example of a failure mode I keep hitting: retrieval returned the right rows and the wrong claim. The agent didn't hallucinate the text — it hallucinated the speech act. A hanging-thread detector that only looks at topic continuity will keep inventing promises. I'd gate that suggestion behind an explicit marker (you literally wrote "coming soon" / "I'll follow up") before it counts as unfinished business.

Collapse
 
eugeniya_ivanova_4a58eadc profile image
Eugeniya Ivanova •

"It hallucinated the speech act, not the text" is a sharper name than anything in my post — I was calling it "inference stated as fact," but yours is more precise, because the words really were mine. Retrieval was correct. What it invented was that those words were a promise to anyone.

And gating on an explicit marker is the better fix. My rule makes the agent ask instead of assert, which is a softer failure but still a failure — it'll still surface a non-promise and make me say no every time. Requiring "coming soon" / "I'll follow up" before a thread counts as unfinished raises the bar at the source instead of catching it at the confirmation. I'm adding that — topic continuity alone clearly isn't enough signal to call something a promise.

Collapse
 
jo-do profile image
Jo Do •

The part that lands is that the promise was real to the reader even though you never made it. Your posting history is a record other parties, human or agent, will mine for commitments, and the writer's own memory is the worst place to look them up. The agent didn't hallucinate the promise; it inferred one from a pattern you'd stopped noticing. That's the uncomfortable thing about giving agents read access to your own record: they end up being more consistent readers of you than you are.

Collapse
 
eugeniya_ivanova_4a58eadc profile image
Eugeniya Ivanova •

"More consistent readers of you than you are" is the line I'll be thinking about. You're right that the promise was real to the reader — the agent didn't misread the data, it read it more literally than the audience ever would, and both of those are truer than my own memory of what I posted.

The uncomfortable part cuts both ways, though. That same consistency is exactly why the read access is useful — it catches the thread I genuinely dropped, the follow-up I actually owe. The failure isn't that it reads me too well. It's that it can't tell the difference between "you committed to this" and "this pattern looks like a commitment," and only one of those is mine to answer for. Which is why the rule ended up being "show me what you saw and let me say what it was," not "decide what I meant." The agent gets to be the consistent reader. It doesn't get to be the author of my intent.

Collapse
 
clarajbennett profile image
Clara Bennett •

Choosing the topic is the part that stalls me too, the actual sentences usually show up once I know which small moment I am telling.