A marketer writes a subject line, runs it past a spam-word checker, gets a clean bill of health, sends it, and watches open rate come in far below expectations anyway. The checker was not wrong. It was just answering a completely different question than "will a human want to open this."
Two separate filters, two separate jobs
A spam filter's job is narrow and mechanical: does this message resemble patterns historically associated with unwanted bulk mail. It looks at specific trigger words, punctuation patterns like stacked exclamation points, ALL CAPS ratios, and sender reputation signals that have nothing to do with the subject line's actual content quality. A line can pass every one of these checks and still be genuinely boring, vague, or misaligned with what the reader actually cares about.
A human reader's decision to open an email is a completely different evaluation, made in under a second, based on whether the subject line signals something specifically relevant and worth the interruption right now. None of the spam filter's checks measure relevance, specificity, or timing, because a spam filter's job was never to judge whether an email is good. It was only ever built to judge whether an email resembles known bad patterns.
Where this creates false confidence
Teams that rely only on a spam-word checker as their subject line quality bar end up with a false sense of security. "It passed the spam check" gets treated as equivalent to "this is a good subject line," when it only actually confirms the much narrower claim "this probably will not get auto-filtered before a human even sees it." Clearing that low bar says nothing about whether the person who does see it in their inbox will actually click.
This is the same category of mistake as treating a document's clean grammar-check result as proof the writing is good. Grammatical correctness and spam-filter cleanliness are both necessary conditions worth checking, and neither one is sufficient on its own.
What actually predicts a human open decision
Specificity is the biggest lever. "Your Q3 report is ready" beats "Important update" because it tells the reader exactly what is inside before they open it, removing the small friction of uncertainty that makes people defer opening something until later, which usually means never.
Timing relevance matters almost as much. A subject line referencing something the reader is actively thinking about right now, a deadline approaching, a question they recently asked, outperforms a generically well-written line with no situational hook, even when both lines are equally clean by spam-filter standards.
Length interacts with device context in a way spam filters do not check at all. A subject line that reads perfectly on desktop can get truncated mid-word on a phone's notification preview, changing its meaning or cutting off the specific detail that made it compelling in the first place. Checking how a line previews across devices catches this; a spam-word scan does not, because truncation has nothing to do with spam patterns.
Running both checks, in the right order
The practical sequence is to write for the human first, then verify the mechanical safety net second. Draft the subject line optimizing for specificity and timing relevance. Then run it through a spam and structural checker, like the one on the EvvyTools site, to catch trigger words, truncation risk, and inbox preview problems you would not have caught by eye. Reversing this order, optimizing for a clean spam score first and hoping relevance follows, tends to produce lines that are technically safe and forgettably generic.
A concrete pair worth comparing side by side
Take two subject lines for the same product update email. "Big news! Check this out now!!" clears most spam filters, since none of those individual words are hard trigger terms, but it tells the reader nothing specific and reads as generic promotional filler. "Your dashboard now shows week-over-week trends" is longer, plainer, and just as clean by spam-filter standards, but it tells the reader precisely what changed and why it might matter to them today. A spam checker would score both lines similarly. Open rate data on lines like these, in practice, consistently favors the specific one, sometimes by a wide margin, because specificity is exactly the thing a mechanical filter has no way to measure.
This is worth testing on your own list if you have any doubt about it. Pick a genuinely vague-but-clean line and a genuinely specific-but-plain line for two comparable sends, and look at which one a real audience actually opens more. In most lists I have seen data from, the specific plain line wins, sometimes by a wide margin, which runs directly counter to the instinct that a punchier, more attention-grabbing line should perform better.
What a combined check actually catches that neither one alone would
Running both checks together surfaces a specific and common failure mode: a subject line that is both specific and clean by spam standards, but that gets truncated on mobile in a way that strips out the exact detail making it specific in the first place. "Your refund of $340 was processed today" reads clearly on desktop, but a mobile preview cutting it at "Your refund of $340 was proc..." loses the reassurance that the process actually completed, which is the entire point of the message. A spam checker sees nothing wrong here, since none of the words are flagged, and a human reading the full line on desktop sees nothing wrong either. Only checking the mobile preview specifically catches this, which is exactly why the structural pass and the human-relevance pass need to happen together rather than as substitutes for each other.
Why marketers keep over-indexing on the wrong filter anyway
Spam-word checkers are easy to build a habit around because they give a clear pass or fail signal in seconds, with no ambiguity. Relevance and specificity are harder to check mechanically, so they get skipped under time pressure even by teams that know, in principle, that they matter more. The fix is not eliminating the spam check, it is a genuinely useful five-second step, but treating it as the last gate rather than the main event, with the bulk of drafting time spent on making the line specific and timely instead of merely mechanically clean.
A quick self-check before you send
Ask whether the subject line would make sense as a text message from someone the reader actually knows, referencing something specific enough that the reader would recognize the topic instantly. If it reads more like a headline written to avoid triggering software than like something a person would actually say, it has probably been over-optimized for the wrong filter.
For a broader look at the gap between what an automated score measures and what actually determines whether content performs, see does Google penalize AI content, and what actually matters.
Further reading: CAN-SPAM Act guidance from the FTC covers what actually constitutes a spam violation legally, Litmus publishes ongoing research on inbox preview behavior across major email clients, and Mailchimp's resource library documents subject line testing practices drawn from aggregate send data across many senders.
Top comments (0)