I run a one-person business in Korea, and a lot of the daily work is done by agents and scheduled scripts: drafting posts, publishing them on a schedule, replying, pulling numbers. The failures that cost me the most were never the loud ones. A crash shows up in a log and you fix it. The expensive ones look exactly like success, or like nothing at all.
Here are five from my own setup, and the check each one gets now.
1. The approvals that never reached the database
A bot sends every draft to my phone with Approve / Edit / Reject buttons, and every tap is supposed to land in a feedback table that the next draft reads first. For a while, approvals weren't being stored. The bot logged the tap, the table never got the row, and nothing complained. The agent just kept "learning" from edits and rejections only.
Check now: one query, run daily. If I tapped approve ten times today and this says zero, the loop is broken no matter how good everything else looks.
select rating, count(*)
from agent_learning.feedback
where created_at > now() - interval '1 day'
group by rating;
2. The scheduled job that never ran once
I installed a posting queue on my VPS in the morning: a systemd timer that runs a Node script every 10 minutes, and the script drives a headless browser with Playwright. I tested it by hand, it worked, I moved on.
By the afternoon, nothing from the queue had been published. The server has two Node binaries. My shell picks up the newer one (22). The service file said ExecStart=/usr/bin/node ..., which is the distro's Node 18, and Playwright refuses to run on it. Every single run failed, quietly, for hours.
Check now:
- Use the full path to the binary you actually tested with (
/usr/local/bin/nodehere), not whatever/usr/binhappens to have. - After installing any timer, run it once through systemd itself and read the journal. A manual run as a different user with a different
PATHproves nothing about the service.
sudo systemctl start my-queue.service
journalctl -u my-queue.service -n 30 --no-pager
- Count outputs, not runs. "The timer fired 30 times" is not the same as "30 things got posted".
3. The upload that stopped at 99%
I swapped the PDF on one of my product pages using the store's newer content editor. The progress bar went to 99% and sat there. The page still looked fine and still had a download button, which pointed at the old file.
I reverted and used the older upload form, which worked. The part that matters: if I hadn't been watching the bar, I'd have announced an update that buyers never got.
Check now: after any product change, open the page logged out and download the file the way a buyer would. Compare the file size or the first page with what you meant to ship.
4. The replies that "weren't there"
I had an agent reply to a few posts on a social platform. Afterwards, the replies didn't show on the post pages, so it looked like the replies had failed, and the obvious next move would have been to post them again. They had gone through. That platform just doesn't always show a fresh reply on the post page right away, but it does show up on the account's own replies tab.
Check now: verify an action where the platform reliably shows it (your own profile or activity tab), and never retry a post just because you can't see it. A duplicate reply is worse than a slow one.
5. The channel that got zero views
I had a set of scheduled posts going into topic communities on one network. They "worked": every one returned an ID, every one was live. I pulled the impression numbers for the first two: 2 and 0.
Nothing was broken, which is why it's on this list. An automation can run perfectly and still be pointless, and you'll only see it if you measure the output that matters, not the status code.
Check now: every channel gets a number I can pull (views, clicks, replies). If the same method gets nothing twice, it stops, and the remaining scheduled posts move out instead of running on autopilot.
The pattern
All five had the same shape: the system reported success at a step before the one I cared about. Tap logged, not stored. Timer fired, not posted. Upload started, not finished. Reply sent, not checked where it shows. Post published, not seen.
So the rule I keep in my agents' instructions now is short: verify the end result where it actually lives, not the step before it. And when something breaks, fix the cause so it can't happen again, not just this one run.
Links
- The feedback loop from #1, running in your browser (Postgres + pgvector in the tab, nothing to install): https://ssap-pa.github.io/self-learning-agent-setup/
- Repo with the schema and a one-minute demo: https://github.com/ssap-pa/self-learning-agent-setup
- The complete version (Python and TypeScript library + CLI, Telegram approval bot that stores every tap and tells you when it can't, a Claude Code plugin on your own Postgres, 28 tests): https://payhip.com/b/grc25
- The playbook this came from, Just Say "Do It" (151 pages, PDF + EPUB), is pay what you want until Sunday, Oct 4 ($0 is fine): https://payhip.com/b/xmZvu. If you read it, an honest review on that page helps a lot.
Top comments (0)