Every social platform rewards the same thing, and it isn't a secret: showing up regularly. The hard part isn't writing. It's deciding what to say today, and on plenty of days there either isn't an answer or you don't feel like hunting for one.
Someone I was talking to on LinkedIn last week told me they stick with their scheduling tool mostly because of the post templates. When they run out of ideas, the templates give them somewhere to start. That surprised me. I would have guessed people got stuck on the writing, not on choosing the topic.
Asking an agent doesn't really fix that either, at least not in the usual setup. You say "write me a post," it gives you one draft based on its first guess, and if you're not convinced, there's often nothing specific to react to. You end up rewriting it yourself, which rather defeats the point.
So this week I built a small skill around one very narrow idea: don't give me a draft first. Give me three angles and let me pick. Three options I can answer with a number.
It works now. The first real run didn't. It failed three times before it gave me anything useful, and all three failures came from the same place: the agent was reading my own data and treating its interpretation as fact.
It read my test posts as things my audience saw
The skill starts by calling list_posts and reading the last twenty. It's looking for two things: topics I've already covered, so it doesn't suggest the same thing again, and threads I've left hanging. A launch with no follow-up. A question I asked and never came back to. Those are often more useful than inventing a new topic because the context already exists and I usually have the material.
It found one. Two posts on my LinkedIn from late July said Publora was coming to Zapier. Nothing after that. So it suggested: finish the thread, you promised this a month ago and never said how it turned out.
Except I hadn't promised anything.
Those posts were Zapier review artifacts. To submit an app there, you have to run every trigger and action inside a live Zap, switch it on, and leave a successful run in the history. That means my account contains posts with names like Zapier validation — update target. They look like normal posts in the same list, and the agent had no idea they weren't meant for an audience.
I added a filter: short posts, near-duplicates minutes apart, anything containing test or validation. But the more useful fix was a rule:
Never state an inference from history as a fact.
"You promised X and never followed up" sounds confident, but it's still an inference from a list that includes QA junk. The agent should say what it found and ask whether it understood it correctly. Once I corrected it, the whole session changed direction in one message. If I hadn't noticed, I could easily have published a post referring to a promise nobody ever saw.
It could only see what went through the product
Once we had the right angle, the skill asked for the one fact it couldn't know: what actually happened with Zapier. I told it we were live but still in beta.
What it didn't ask was whether I'd already written about it somewhere else. I had. That morning I'd published a four-minute Dev.to article with the eight rounds of review, the REST Hooks, and the JavaScript I ended up putting in a field labeled "label". None of that existed in Publora, so the skill had no way to know about it.
That limitation is obvious in theory: an agent only knows the data you give it. In practice, it's easy to forget until it starts repeating something you've already published elsewhere.
So now, whenever it asks for a missing fact, it also asks whether I've already written about the topic somewhere else: article, changelog, release notes, whatever.
The slop checker passed a draft that was still slop
I run every outgoing text through a script that flags machine-sounding writing: vocabulary, repeated structures, the usual tells that become easy to spot once you've collected enough of them.
The draft passed.
I showed it to a colleague anyway, and she said she could still see AI slop in it. She was right.
Three of the strongest lines had been lifted almost directly from my Dev.to article from that morning. "The requirement I reread three times, certain I'd misunderstood." Fine line on its own. Less fine when the same person has already seen it six hours earlier. No vocabulary checker is going to catch that.
The other problems were structural. A withheld hook: "with a Beta tag, which I'll get to." A neat paradox: "you have to prove the app is being used before it's allowed to exist." A three-item list whose last item worked mostly because of the rhythm. The word "literally". None of those is automatically bad. Together, in a short post, they started sounding very familiar.
The rewrite was 720 characters. The version that passed the script was 1198.
So I kept the checker, but moved it earlier in the process. It catches the things that exist inside the text. What it can't see is everything around the text: what I published yesterday, who is reading, or whether I've stacked too many individually reasonable choices into something that sounds generated.
I'd still put a checker somewhere between draft and publish. It costs very little once it exists, catches the obvious layer, and doesn't get tired at six in the evening. Build one, borrow one, or use a checklist. Just don't mistake it for the whole review.
What it looks like now
There are eight categories and forty angles. Every angle has a core: one fact that has to come from the person using the skill, whether that's the mistake, the number, or the tool name. The instructions repeat this in several places because inventing that core is the easiest way to produce something plausible and false under someone's name.
There are also four reference files behind the angles: what the first two lines need to do before the feed collapses the rest, how to infer voice from existing posts instead of asking someone to describe their own voice, the tells a script can't catch, and how the same angle changes between LinkedIn and X.
It stops at the draft. Nothing publishes automatically.
It's MIT, it works on its own, and connecting Publora lets it read your feed and schedule the result: publora-team/publora-post-ideas
For the record, I'm not an engineer. My job is getting people to use our product, and I built this with Claude: it wrote the code, I led, tested, and sent it back when it got things wrong. Most of the useful work ended up being in those three failures, because each one showed me a rule I hadn't written yet.
Has an agent ever been confidently wrong about your own data? I'm curious what it misunderstood, and what you changed after that.
Top comments (11)
Really interesting failure sequence, Eugeniya. I think all three cases share the same deeper problem: the evidence was real, but the meaning assigned to it was stronger than the evidence justified. 🔍
The Zapier posts are a great example. The agent correctly observed that those records existed, but “you published this to your audience” and “you promised a follow-up” were additional claims that required provenance the post history did not contain.
So I really like the rule:
observation != interpretation != fact
I would almost want every historical item the agent consumes to carry context such as source, audience, purpose and confidence before it is allowed to reason from it.
The cross-platform case is the same problem from the opposite direction. The agent did not hallucinate anything, it simply had an incomplete evidence boundary and no way to know that the missing context existed elsewhere.
And the slop checker is probably my favorite failure here. It passed because it was measuring properties inside the draft, while the thing your colleague noticed depended on context outside the draft: reuse, recent history and the cumulative effect of individually reasonable patterns.
So the checker was not necessarily wrong. The property being asked of it was larger than its observation surface.
That makes the final “stop at the draft” boundary especially important. The model can propose, the checker can catch a subset of problems, but neither gets publishing authority simply because both are green.
I think the architecture starts looking like:
history -> provenance/context -> inference -> human confirmation -> draft -> local checks -> publish decision
Really good example of why agent memory becomes much more useful when it remembers where a fact came from and what it actually proves, not just the fact itself. 🧠🔐
"The meaning assigned to it was stronger than the evidence justified" — that's the sentence I wish I'd opened the article with. All three were three stories to me; you've written them as one, and you're right.
The provenance idea is where it gets practically interesting. Attaching source, audience, purpose, confidence to each historical item is exactly what would have stopped the Zapier misread at the root — a QA run tagged audience: none, purpose: test never gets read as a promise. The catch is that provenance has to be captured at write time, by the thing that created the record, and most systems don't. My test posts went into the feed with no marker because nothing along the way had a reason to add one. So the architecture is right, but it pushes the cost upstream: the feed has to start carrying context it currently throws away, and retrofitting that onto records that already exist is the hard part.
Your read on the slop checker is more generous than mine and also more correct. "The property being asked of it was larger than its observation surface" is the honest version of what I called a failure. It answered the question it could see. I was holding it responsible for a question that lived outside the file, which isn't the checker's fault, it's a category error on my part about what that tool is for.
And the pipeline you drew is the one I backed into without naming: history → provenance → inference → human confirmation → draft → local checks → publish. The one thing I'd underline is that the human-confirmation step and the publish-decision step are different gates doing different jobs — one asks "is the premise real," the other asks "should this go out," and collapsing them is how "both checks are green" quietly becomes publishing authority. Neither green light is the same as a decision.
This is a fascinating breakdown, especially the part about structural "slop". You're totally right that vocabulary checkers are fighting yesterday's war. The real AI tells are the predictable rhythms: the dramatic single-sentence paragraph, the neat paradoxes, and those perfectly balanced three-item lists.
Since you built the logic with Claude, I'm curious about how you are handling the prompt for the actual draft phase now. Have you tried passing negative constraints to specifically ban those structural clichés (e.g., "Do not use three-item lists" or "Avoid dramatic one-line paragraphs"), or does Claude end up ignoring them when trying to mimic your voice? Would love to know if negative prompting worked for your use case!
Great question, and the honest answer is that negative constraints in the draft prompt didn't carry the load for me, so I stopped leaning on them. "Don't use three-item lists" works right up until the model is also trying to match my voice and hit a length — then the structural habit slips back in, because it's not reaching for a list on purpose, it's falling into a rhythm, and a rule it isn't actively thinking about doesn't fire.
So I moved the anti-slop work out of the draft prompt and into a separate pass that runs after. Generate first, then check against the tells as its own step, rather than asking the model to avoid them while it's busy doing three other things. The catch — which is basically the whole point of that article — is that a checker only sees what's inside the text, so the structural pass still misses the cross-post repetition and the "four defensible choices add up to a machine" problem. Those I still catch by reading.
Short version: negative prompting helped a little, a post-draft check helped more, and neither one replaces a human read. Have you had better luck getting the constraints to actually stick in-prompt?
That makes perfect sense. I've had similar struggles where negative constraints work for a sentence or two, and then the model completely forgets them by paragraph two! Treating the drafting and the editing as two distinct steps instead of one mega-prompt is a brilliant workaround. Thanks for sharing the behind-the-scenes on this, it’s definitely given me some ideas for my own workflows!
The Zapier validation posts are a clean example of a failure mode I keep hitting: retrieval returned the right rows and the wrong claim. The agent didn't hallucinate the text — it hallucinated the speech act. A hanging-thread detector that only looks at topic continuity will keep inventing promises. I'd gate that suggestion behind an explicit marker (you literally wrote "coming soon" / "I'll follow up") before it counts as unfinished business.
"It hallucinated the speech act, not the text" is a sharper name than anything in my post — I was calling it "inference stated as fact," but yours is more precise, because the words really were mine. Retrieval was correct. What it invented was that those words were a promise to anyone.
And gating on an explicit marker is the better fix. My rule makes the agent ask instead of assert, which is a softer failure but still a failure — it'll still surface a non-promise and make me say no every time. Requiring "coming soon" / "I'll follow up" before a thread counts as unfinished raises the bar at the source instead of catching it at the confirmation. I'm adding that — topic continuity alone clearly isn't enough signal to call something a promise.
The part that lands is that the promise was real to the reader even though you never made it. Your posting history is a record other parties, human or agent, will mine for commitments, and the writer's own memory is the worst place to look them up. The agent didn't hallucinate the promise; it inferred one from a pattern you'd stopped noticing. That's the uncomfortable thing about giving agents read access to your own record: they end up being more consistent readers of you than you are.
"More consistent readers of you than you are" is the line I'll be thinking about. You're right that the promise was real to the reader — the agent didn't misread the data, it read it more literally than the audience ever would, and both of those are truer than my own memory of what I posted.
The uncomfortable part cuts both ways, though. That same consistency is exactly why the read access is useful — it catches the thread I genuinely dropped, the follow-up I actually owe. The failure isn't that it reads me too well. It's that it can't tell the difference between "you committed to this" and "this pattern looks like a commitment," and only one of those is mine to answer for. Which is why the rule ended up being "show me what you saw and let me say what it was," not "decide what I meant." The agent gets to be the consistent reader. It doesn't get to be the author of my intent.
Choosing the topic is the part that stalls me too, the actual sentences usually show up once I know which small moment I am telling.
Some comments may only be visible to logged-in visitors. Sign in to view all comments.