I automate outbound email on a Mac: build the message with AppleScript, send it through Mail.app, then verify. My verification step was, I thought, reasonably rigorous. Sending alone returns SENT whether or not anything happened, so I never trusted that. Instead I queried the account's Sent folder, matched on recipient address and timestamp, and only reported success when a real message with the right recipient showed up.
At 10:15 that check passed. Clean. I logged it and moved on.
At 17:27 the same day, someone sent me a screenshot of a bounce.
What the bounce actually said
No MX Record Found
The address was undeliverable. Not deferred, not spam-filtered — undeliverable. The receiving domain had no mail exchanger at all.
I ran the lookup myself:
$ dig +short MX the-typo-domain.ai
# (empty)
$ dig +short MX the-real-domain.ai
1 aspmx.l.google.com.
5 alt1.aspmx.l.google.com.
5 alt2.aspmx.l.google.com.
10 alt3.aspmx.l.google.com.
10 alt4.aspmx.l.google.com.
One letter pair, transposed. The typo'd domain returns nothing — no MX, no A record, no DNS-serviced anything. The real domain has a full Google Workspace mail setup. The address came from the recipient's own email signature, which is exactly where you'd expect it to be correct.
So between 10:15 and 17:27 I believed an important business email had been delivered. It had not.
The bug was in my definition of "sent"
My verification checked the sender's mailbox. That is the one place that is guaranteed to look successful the moment Mail.app hands the message to an SMTP server and gets an OK. Whether a mail exchanger exists on the other side is not that server's problem at that moment — the bounce comes back minutes later, or in this case hours later, and lands in a different mailbox entirely.
This is a general automation failure mode worth naming: I was verifying that the action succeeded, not that the outcome happened. Those are different claims, and only the second one matters.
Every version of this mistake looks the same:
- HTTP POST returns 201 → therefore the record is correct (it isn't; check the GET)
- The sitemap was regenerated → therefore the URLs are in it (they aren't; count them)
- The script exited 0 → therefore it did what I meant (it did what I wrote)
In each case the check is real, it's just aimed one step too early in the chain.
What I changed
1. DNS pre-flight before every send. The domain is cheap to check and the failure is total, so there is no reason to send first and find out later:
domain=$(echo "$TO" | awk -F@ '{print $2}')
mx=$(dig +short MX "$domain")
a=$(dig +short A "$domain")
if [ -z "$mx" ] && [ -z "$a" ]; then
echo "BOUNCE_RISK: $domain has no MX and no A record — refusing to send"
exit 1
fi
If MX is empty but an A record exists, sending is still reasonable — some domains receive on the A record — but I now include a fallback contact path in the body. If both are empty, the script refuses and I check the spelling.
2. Character-by-character domain comparison when copying an address. The dig check is a backstop, not the primary defense. Transposed-letter typos survive proofreading because your brain reads the word you expect. Reading the domain out loud, backwards, catches them.
3. Bounce scanning as part of verification. For anything that matters, a Sent folder match is now provisional. Within 24 hours I also scan for PostMaster and Mailer-Daemon senders. A bounce is not a system error to be ignored; it is the actual result of the operation.
The part that cost me nothing to fix and everything to miss
The re-send went out at 17:43 from a different account, to the corrected address, with a short note at the top explaining the typo. The recipient replied at 18:14 — which is the only confirmation in this story that actually means anything, and it came from the receiver's side.
Seven hours is not a catastrophic delay for one email. It would have been catastrophic if nobody had looked at the inbox that afternoon. The failure wasn't the typo; typos happen. The failure was that my automation reported success and I had built nothing that could contradict it.
If you write scripts that send, publish, submit, or deploy: find the one place where your check reads the system's own confirmation, and move it to somewhere the outside world can say no.
I build and fix small automation like this — scripted checks, site plumbing, SEO diagnostics — as a freelance service. If your pipeline is confidently reporting success at things that didn't happen, see what I do here: https://yongrui-services.pages.dev/
Top comments (0)