DEV Community

ssapable
ssapable

Posted on Fully Autonomous

My daily marketing job came out ~35x cheaper as a script with one AI call than as a live agent

I run a one-person online course business in Korea, and AI agents do a lot of the daily marketing: posting, replying to people who show up, pulling numbers. The first time I automated a daily job, I let the agent do the whole thing live. It worked. It was also the most expensive way to do it.

Here's what I'd tell anyone automating marketing for a small business, in the order I learned it.

1. Pick how the agent touches the outside world

Every automation reaches a service in one of five ways, and picking the wrong one is how you burn tokens and get things that break weekly.

Way What it is Use it when
MCP A standard plug that lets agents use a tool (Claude calls them connectors) The service offers an MCP server. Easiest route
API The company's official window for software No MCP, but there's an official API. Accurate, good for bulk work, needs a key and code
Browser / computer use The agent clicks and types like a person No API at all. Browser automation reads the page structure and is cheaper; computer use works on desktop apps but is slower and eats tokens
Command line git, npm, ffmpeg… typed by the agent Files and programs on your own machine
Code execution The agent writes a script and runs it No existing tool does the job. My video-editing helper is a Python file an agent wrote

A sixth piece decides when things run: a scheduler (cron, Windows Task Scheduler, systemd timers). It does nothing on its own; it just rings and starts one of the five.

My rule of thumb: MCP if it exists, then the official API, then the browser. Screen-based methods break when a layout changes, so they're the last resort.

2. Look at what a "daily marketing job" actually is

My first job was simple: every evening, open the feed, look at ten posts, write a reply on the best one, post it. Written out as steps:

  1. Open the feed
  2. Collect the posts
  3. Rank them by engagement
  4. Write the reply ← the only step that needs a model
  5. Check it
  6. Post it

Five of the six steps are identical every day. When the agent does all six live ("agent mode"), it reasons about every click, and the conversation keeps growing, so every step re-reads everything before it.

I had the agent estimate both versions of the same job at API token prices (I use subscriptions, so this was a comparison, not a bill). Agent mode came out at a few dollars per run. A script that does steps 1-3 and 5-6 in code and calls a model once for step 4 came out at a few cents. About 35x in my test. Once, it doesn't matter. Every day for months, it decides whether you keep doing it.

3. Ask for the split explicitly

Agents won't always split it on their own. The one sentence that changed my plans:

When you design this, put repetitive logic and steps in code or scripts.
Use an AI model only for steps that need judgment or content creation.
Enter fullscreen mode Exit fullscreen mode

I also start in plan mode and read the plan before approving it, including the English terms I don't know yet ("dry run" means a test that doesn't actually post). Plans built on old work get the line "Ignore my earlier plans. Re-plan from scratch using only what I just wrote."

4. Choose the model for the one step that matters

My first automated reply was a stiff, generic compliment, the kind anyone spots as AI in a second. I isolated the cause instead of shrugging:

  • The instructions had drifted. When the agent turned my writing rules into a reusable skill, it had lightly reworded them. I re-applied the originals word for word.
  • The model mattered too. I ran the same prompt on three models. The lightest one gave broad praise and stock phrases. The mid-tier one summarized the post and added a feeling. The top model at a low reasoning setting wrote short, specific replies, so that's the one the script calls.

Because the model runs once per job instead of on every click, using the better model is affordable.

5. Put a human on anything customer-facing

Replies to real people go to my phone first with Approve / Edit / Reject. Every tap is stored, and the next draft reads the closest past corrections before it writes. That's a separate post (how the feedback loop works), but the short version: automation decides when and what to draft; I still decide what goes out under my name.

6. Move it off your laptop

I started with Windows Task Scheduler (7:11 p.m. daily, after people get off work). It's the easiest place to begin because your logged-in browser sessions are already there. The catch: when the computer is off, nothing runs.

Now the scheduled jobs live on a small VPS. Posting queues for X and Threads run every 10 minutes from systemd timers, and they post whatever is due. My laptop can be closed and the marketing still goes out.

7. Fix the cause, not the run

The next scheduled run failed with a browser connection error. Instead of patching it by hand, I had a read-only review agent read the logs and find the root cause, then had the main agent fix it, commit, and test on schedule. Then I wrote this into my global AGENTS.md, and it's the rule I'd keep if I could keep only one:

When a problem occurs, don't stop at a temporary or manual fix. Find the root cause so it doesn't happen again, propose a solution, change your approach or move up to a stronger model if needed, and keep going until the final goal is reached.

Keep AGENTS.md short, though. Every session reads it first.

Checklist

  • [ ] Wrote the job out as steps and marked the ones that need judgment
  • [ ] MCP or official API before browser automation
  • [ ] Repetitive steps in code, one model call where it matters
  • [ ] Compared models on the same prompt for the writing step
  • [ ] Approval step for anything customer-facing
  • [ ] Scheduler on an always-on machine
  • [ ] Root-cause rule in AGENTS.md

Links

Top comments (2)

Collapse
 
brianainews profile image
Brian · AI News •

That 35x split is the part most agent demos skip. Keep repetitive work deterministic and spend model calls where judgment actually matters. I covered a related test of AI video agents and avatars here youtu.be/0YPSv4keATc

Collapse
 
deanlee profile image
Dean Lee •

The 35x gap isolates where agent architecture actually burns capital: running deterministic orchestration through an autoregressive model.

When a pipeline relies on an LLM to navigate DOM trees, parse JSON responses, and manage control flow, every intermediate turn serializes the prior state back into the prompt. That turns fixed operational steps into quadratic token drag. The execution cost scales with the history of the run rather than the complexity of the judgment being made.

Treating the model strictly as an isolated pure function (passing a bounded prompt and receiving a single completion) shifts state management back to standard runtime logic where execution is free. The economic ceiling for production automation has never been reasoning capability; it is whether the control plane can keep deterministic state out of the context window.