DEV Community

StareBrain
StareBrain

Posted on

Email Validation Has Three Right Answers, Not Two

Shipped a small fix today that's worth writing up because the obvious version of it is wrong, and the wrong version is what almost every email-validation snippet online actually does.

The setup

A reader pointed out our waitlist form only checked that an email was shaped like an email — matched a regex, had an @ and a dot. It never checked whether the domain behind it could actually receive mail. someone@thisdomaindoesnotexist.com sailed straight through.

The obvious fix is an MX lookup: resolve the domain's mail-exchange records, reject if there aren't any. That's a five-minute change. Here's the five-minute version:

const records = await dns.resolveMx(domain);
if (!records || records.length === 0) {
  return reject("No mail server found for this domain.");
}
Enter fullscreen mode Exit fullscreen mode

This is wrong, in a way that won't show up in testing and will show up in production.

What's actually wrong with it

A DNS lookup has three possible outcomes, not two. It can succeed and find records (domain's good). It can succeed and find no records (domain can't receive mail, genuinely). Or it can fail to get a clear answer at all — a timeout, a SERVFAIL, a resolver hiccup, a slow network path to whatever DNS server is being queried.

The five-minute version above only has two branches: records exist, or they don't. A timeout throws, lands in neither branch explicitly, and depending on how the catch is written, usually gets treated as "no records" by default, because that's what the catch block does with anything it didn't expect. A real, deliverable Gmail address, on a day when the resolver is slow, gets rejected as if the domain doesn't exist.

That's the bug. Not "the check doesn't work," but "the check conflates two things that aren't the same: the domain has no mail server and we couldn't find out whether it does."

What shipped instead

type MxResult = "has_mx" | "no_mx" | "unchecked";

async function checkMx(domain: string): Promise<MxResult> {
  const timeout = new Promise<"unchecked">((resolve) =>
    setTimeout(() => resolve("unchecked"), MX_TIMEOUT_MS)
  );

  const lookup = (async (): Promise<MxResult> => {
    try {
      const records = await dns.resolveMx(domain);
      return records && records.length > 0 ? "has_mx" : "no_mx";
    } catch (err: any) {
      if (err?.code === "ENOTFOUND" || err?.code === "ENODATA") {
        return "no_mx"; // definitive: this domain has no mail server
      }
      return "unchecked"; // inconclusive: SERVFAIL, resolver timeout, etc.
    }
  })();

  return Promise.race([lookup, timeout]);
}
Enter fullscreen mode Exit fullscreen mode

Three states, mapped to three different outcomes at the call site: no_mx rejects the signup outright, with a message about checking for a typo. has_mx proceeds normally. unchecked also proceeds — rejecting a real signup because our DNS resolver had a slow five seconds is a worse failure than letting through an occasional typo — but it gets written to the record as mxStatus: "unchecked", not silently merged into "verified."

That last part is the actual point of the whole change. It would have been easy to write unchecked and has_mx as the same branch, since they lead to the same immediate action (let the signup through). But they're not the same fact, and collapsing them loses something real: if DNS lookups start timing out at scale, or a particular registrar's domains always come back inconclusive, that's invisible unless "unchecked" stays its own labeled state instead of disappearing into "verified" the moment it stops blocking anything.

The general shape

This is the same mistake as writing if (success) { ... } else { ... } for any operation that can fail in more than one way. The binary branch is almost always built around what action to take, not around what's actually true. Those often need the same action (let the signup through) while being different facts worth keeping separate (confirmed good vs. we don't actually know).

The cheap test: does your error handling ever write down "we don't know," or does every code path resolve to something that looks like a definite answer? If it's the latter, you're probably doing the same thing we were doing yesterday — treating an inconclusive result as a known one because the code needed to decide what to do next, and deciding and knowing got merged into one step.

Top comments (0)