DEV Community

Craig Solomon
Craig Solomon

Posted on

How an auto-publishing agent earns its permissions, channel by channel

A publish permission wants to be a boolean. AUTO_PUBLISH=true, ship it, go do something else.

The trouble with the boolean is what it encodes: that the agent is trustworthy in general. Trust is not general. A drafting agent can be perfectly sound writing a technical post for a dev feed and be a liability in a subreddit where a link in the body reads as spam. Same model, same prompt, different blast radius.

So in Content Agent Pro the auto-publish check is not a flag I set. It is a conjunction evaluated per (product, channel) pair, and most of the terms can go false without me touching the agent at all.

The gate

def can_autopublish(pair, cfg):
 if not cfg.master_switch:
 return False, "master auto-publish switch is off"
 if not pair.channel.free_api:
 return False, "channel has no free publishing API"
 if not pair.channel.credentials_present:
 return False, "credentials missing for this channel"

 reviewed = pair.approved + pair.rejected
 if reviewed < MIN_REVIEWED: # 10
 return False, f"only {reviewed} reviewed drafts"
 if pair.approved / reviewed < MIN_RATE: # 0.90
 return False, "approval rate below threshold"

 return True, "graduated"
Enter fullscreen mode Exit fullscreen mode

MIN_REVIEWED = 10 and MIN_RATE = 0.90 are the graduation thresholds. Everything interesting about this function is in the details around those two numbers.

Evidence is reviewed drafts, not drafts

reviewed is approved + rejected. Pending drafts are not in the denominator, and they should not be. A draft nobody looked at is not evidence about anything. If I let total drafts drive the gate, a quiet week where the scheduler runs and I never open the dashboard would push a pair toward autonomy on the strength of work no human ever read. That is the exact inversion of what the gate is for.

This is worth checking in any track-record gate you build. The metric has to be judged outcomes, and the judgment has to have actually happened. Anything that accrues on its own is a clock, not a record.

The unit is the pair, not the agent

State lives per (product, channel). Product A on Dev.to can be graduated while product A on Reddit is still fully manual and product B on Dev.to has never been reviewed.

This falls out of how the drafting works. Each channel definition carries its own voice rules and link discipline, and link discipline in particular varies hard: footer link, profile-only, comment-only, no link at all. Nine channel definitions ship with the kit, and a draft that is honest and well-formed for one of them can be wrong for the next. The failure modes are not shared, so the trust record should not be either.

Practically, that means your schema key is a composite, and your approval counters are columns on that pair row rather than globals. If you find yourself writing SELECT avg(approved) FROM drafts, you have built a reputation score for the model. What you want is a reputation score for a specific job.

Two terms that are not about performance at all

master_switch and credentials_present are not evidence, and they are in the conjunction on purpose.

The master switch is the thing I can flip to make every graduated pair manual again, in one place, without editing per-pair state. If the only way to stop an agent is to walk through its permission records and revoke them one by one, you do not have a stop.

free_api is a constraint I encoded because of what the failure looks like when it is absent. Auto-publishing is only allowed on channels whose publishing API is free. Cost changes the shape of a runaway: a loop that posts too often on a free endpoint is an embarrassment you delete, and the same loop against a metered endpoint is a bill. The kit ships publishers for Dev.to and Bluesky, plus a Hashnode publisher that ships disabled, because Hashnode ended free API access and it now works only on a Hashnode Pro plan. That disabled publisher is the constraint doing its job in public rather than a feature I forgot to finish.

The rate only climbs if rejection carries information

A 90% approval threshold is only reachable if rejections change the next draft. So rejection in the dashboard is reject-with-reason, and the reason goes into a bounded window that gets injected into the next prompt for that pair: the last 8 reject reasons and the last 6 titles for the product and channel.

Both windows are bounded, for different reasons. Reject reasons go stale as the voice settles, and an unbounded list of them turns the prompt into an archive of complaints the model has already absorbed. Recent titles are there so the next draft does not restate the last one under a new headline.

If you are building this part yourself, the thing to get right is that the reason field is free text written at the moment of rejection. A dropdown of categories is easier to aggregate and far weaker as an input, because the useful part is the sentence, not the label.

A gate is not a substitute for a check

Here is the part that took a design decision rather than a threshold. A track record is a statement about the past. Approval rate cannot catch a claim the model invents today.

So the gate sits on top of deterministic validators that run on every draft, graduated or not. They are code, not prompt instructions. Any dollar amount, percentage, or large number that does not appear in the product's fact sheet hard-fails the draft. Hype words fail. AI-tell phrases fail. Unfilled placeholders fail. Per-product earnings language fails. Em dashes are handled differently: a sanitizer strips them, and the draft is not retried for it, because a character-level fix does not need a model round trip.

The division of labor is: validators define what is never allowed to publish, and the gate defines when a human stops having to look. Those are different questions, and a system that only has one of them either publishes fabrications with a clean track record or asks you to catch fabrications by eye forever.

In the dashboard, every draft card shows a validator flag count and the full findings, and /gates renders the graduation matrix so the state of each pair is visible rather than inferred.

Where this does not help

The honest limits, because they are the reason to decide deliberately rather than copy the pattern.

The threshold is not a guarantee. Ninety percent is not a hundred, and a graduated pair will eventually publish something you would have edited. The validators are what keeps that from being a fabricated number under your name. They do not keep it from being a mediocre post.

The record is historical and the model is not frozen. Change your product's fact sheet or your channel voice rules and the accumulated approval rate describes a job that no longer exists. Resetting a pair's counters after a material change to its inputs is a reasonable thing to add, and it is the kind of thing you have to decide for yourself, because only you know which config changes count as material.

Gates also assume you actually review. They convert human attention into autonomy. If the drafts pile up unread, nothing graduates, and correctly so.

And if you publish to one channel and you write the posts yourself, none of this is worth the schema. The mechanism earns its keep when there are several products, several channels with genuinely different rules, and a real cost to being wrong in public.

Content Agent Pro is the self-hosted, MIT-licensed version of this: a Python agent and a Next.js dashboard over one SQLite file, which I built, maintain, and run myself.

https://fulcrumenterprises.tech/go/content-agent-kit-pro/?c=devto

Top comments (0)