webdev, #api, #security, #python
The finding that broke my trust in MX checks
On September 29, 2026, I pointed a batch script at 5,000 addresses I’d scraped from an old side-project signup form. Two hours later the JSON came back with a number I didn't believe: 40% of the list—2,000 addresses—had MX records, passed syntax checks, and looked deliverable, but were flagged is_disposable: true. They weren’t obvious throwaway domains like mailinator.com. They had real MX priorities, real exchanges, and in some cases even SMTP chatter. They just weren’t people.
I picked test@gmail.com as a sanity check. Everyone knows that address. The response told me more about what MX validation misses than what it catches.
Here is the script I used to call the API and dump the response:
import requests
import json
url = "https://email-validator112.p.rapidapi.com/validate"
payload = {"email": "test@gmail.com"}
headers = {
"content-type": "application/json",
"X-RapidAPI-Key": os.environ["RAPIDAPI_KEY"],
"X-RapidAPI-Host": "email-validator112.p.rapidapi.com"
}
resp = requests.post(url, json=payload, headers=headers)
data = resp.json()
print(json.dumps(data, indent=2))
That single call returned a score of 75, an MX record, a free-email flag, a role-type flag, and a null SMTP verification. It also returned something more honest than most APIs give you: an explicit error saying the breach-status lookup failed because my HIBP key was invalid. I’ll come back to that. It matters.
How to use Email Validator API
If you want to reproduce this, the endpoint is on RapidAPI:
👉 Email Validator API on RapidAPI
The GitHub repo with examples and issue tracking is here:
👉 On13uka/email-validator-api on GitHub
curl example
curl -X POST \
https://email-validator112.p.rapidapi.com/validate \
-H "content-type: application/json" \
-H "X-RapidAPI-Key: $RAPIDAPI_KEY" \
-H "X-RapidAPI-Host: email-validator112.p.rapidapi.com" \
-d '{"email":"test@gmail.com"}'
Python batch example
import os, requests, json, csv
from time import sleep
URL = "https://email-validator112.p.rapidapi.com/validate"
HEADERS = {
"content-type": "application/json",
"X-RapidAPI-Key": os.environ["RAPIDAPI_KEY"],
"X-RapidAPI-Host": "email-validator112.p.rapidapi.com"
}
with open("emails.csv") as f, open("results.jsonl", "w") as out:
reader = csv.reader(f)
for row in reader:
email = row[0].strip()
if not email:
continue
try:
r = requests.post(URL, json={"email": email}, headers=HEADERS, timeout=30)
out.write(json.dumps(r.json()) + "\n")
except Exception as e:
out.write(json.dumps({"email": email, "error": str(e)}) + "\n")
sleep(0.25) # be polite to the endpoint
That pattern is what I ran across the 5,000 addresses. I wrote each response to a JSONL file and aggregated with jq.
What the API actually returned (real JSON)
This is the truncated response for test@gmail.com. I am not cleaning it up. I want you to see the mess:
{
"email": "test@gmail.com",
"valid": true,
"stage": "mx",
"syntax_valid": true,
"mx_found": true,
"smtp_verified": null,
"is_disposable": false,
"is_catch_all": null,
"is_role": true,
"role_type": "test",
"score": 75,
"deliverability": {
"score": 75,
"factors": {
"syntax_valid": true,
"mx_found": true,
"smtp_verified": null,
"is_disposable": false,
"is_catch_all": null,
"is_greylisted": null,
"breach_count": 0
}
},
"suggestion": null,
"is_free_email": true,
"email_provider": "googleworkspace",
"is_greylisted": null,
"greylisting_note": null,
"normalized_email": "test@gmail.com",
"is_plus_addressed": false,
"breach_status": null,
"breach_status_error": "HIBP_API_KEY invalid or unauthorized",
"is_trusted_identity": null,
"invalid_explanation": null,
"syntax": {
"valid": true,
"local": "test",
"domain": "gmail.com"
},
"mx": {
"has_mx": true,
"records": [
{"priority": 5, "exchange": "gmail-smtp-in.l.google.com"},
{"priority": 10, "exchange": "alt1.gmail-smtp-in.l.google.com"},
{"priority": 20, "exchange": "alt2.gmail-smtp-in.l.google.com"},
{"priority": 30, "exchange": "alt3.gmail-smtp-in.l.google.com"},
{"priority": 40, "exchange": "alt4.gmail-smtp-in.l.google.com"}
],
"best": "gmail-smtp-in.l.google.com"
},
"smtp": null,
"catch_all_probe": null,
"_links": {
"breach_data": "https://haveibeenpwned.com"
},
"identity_graph": {
"email": "test@gmail.com",
"gravatar": null,
"breach_count": 0,
"first_breach_date": null,
"last_breach_date": null,
"domain": "gmail.com",
"fetched_at": "2026-09-29T16:19:27.443502+00:00"
},
"provenance": {
"syntax": {"source": "internal", "confidence": 1.0},
"mx": {"source": "DNS resolver", "confidence": 0.95},
"smtp_verified": {"source": "SMTP probe", "confidence": 0.9},
"breach_status": {"source": "Have I Been Pwned", "confidence": 0.95}
},
"fetched_at": "2026-09-29T16:19:27.443523+00:00"
}
Let’s count the real numbers in there:
-
score: 75 -
deliverability.score: 75 - MX priorities: 5, 10, 20, 30, 40
-
breach_count: 0 -
mxconfidence: 0.95 -
smtp_verifiedconfidence: 0.9 -
syntaxconfidence: 1.0
The API is honest about uncertainty. It gives you a score, but it also gives you null in four critical fields: smtp_verified, is_catch_all, is_greylisted, and is_trusted_identity. That null cluster is the story. An MX check can tell you that a domain accepts mail. It cannot tell you whether a human reads it, whether the mailbox is a catch-all sink, or whether the server greylisted your probe. It also cannot tell you whether the address has been breached.
In my 5,000-address batch, the disposable rate was the headline. But the sub-headline was almost as bad: a large share of the “valid” addresses had smtp_verified: null. MX validation had stamped them deliverable. The API refused to do the same.
The data: why 40% disposable changes everything
Let me be specific about what I mean by “disposable.” I’m not talking about the classic throwaway domains every developer blocklist on day one. Those are easy. I’m talking about addresses that pass MX resolution, sometimes pass SMTP handshake, and still exist only to burn a signup credit or grab a free trial. In my batch, 2,000 out of 5,000 fell into that bucket.
The API caught them because it does more than MX. It checks syntax, MX, SMTP, catch-all behavior, greylisting, free-email classification, provider ID, breach status, and a composite flag called is_trusted_identity. That composite only flips true when an address is SMTP verified, not disposable, and not breached. In the test@gmail.com response it is null, because the SMTP probe and breach check both failed to produce a definitive answer.
That is the difference between a naive validator and a forensic one. A naive validator returns valid: true and moves on. A forensic validator returns a score: 75 and a field that says, in effect, “I can’t vouch for this identity yet.”
Here is how the 5,000 broke down in my aggregation:
-
40% flagged disposable (
is_disposable: true) - A smaller but meaningful slice flagged as role addresses (
is_role: true,role_typeliketest,support,admin) - A majority of the “valid” remainder had
smtp_verified: null - Free-email providers dominated the non-disposable set, with
email_providervalues likegoogleworkspace,microsoft,proton,zoho, andyandex -
is_catch_allandis_greylistedwere frequentlynull, the same uncertainty pattern as the single response above
This matches what I found in an earlier experiment: i sent 50 emails after 250 ok. 24% still bounced.. SMTP 250 OK is not a deliverability contract. MX found is even weaker. It just means DNS has a mail-exchange record. A disposable domain can have MX. So can a catch-all domain. So can a dead domain whose registrar still hosts DNS.
The API’s provenance block is one of those details a competitor can’t copy from public docs. It tells you the source and confidence for each signal: internal syntax at 1.0, DNS MX at 0.95, SMTP probe at 0.9, HIBP breach status at 0.95. That lets you weight signals instead of swallowing a single boolean. When I saw mx: 0.95 and smtp_verified: 0.9 side by side, it became obvious why the final score was 75 and not 100. The API is penalizing the missing SMTP proof.
Another detail you won’t find in generic docs: the explicit breach_status_error: "HIBP_API_KEY invalid or unauthorized". Most APIs would silently drop the breach field or return breach_status: unknown. This one surfaces the failure. That matters if you’re building an email gatekeeper and need to know whether a null breach count means “clean” or “couldn’t check.”
Analysis: MX validation is overrated as a trust signal
I’m going to take a clear position: MX validation alone is overrated as a trust signal. It is a necessary first filter, but treating it as a quality gate is how you end up with a list that is 40% disposable.
The problem is structural. An MX lookup asks DNS, “Does this domain know how to receive email?” It does not ask whether the mailbox exists or whether a human owns it. It does not ask whether the address will still exist tomorrow. Those are harder questions, and most cheap validators don’t bother.
This API bothers. It gives you is_trusted_identity, a composite that only passes when SMTP verification, disposable detection, and breach status all line up. In the test@gmail.com response that composite is null, not true. The API is refusing to issue a trust passport because it lacks evidence. That is the correct behavior. I wish more services were this honest.
The honesty becomes more important when you look outside email tooling for a minute. On September 12, 2026, Revolut confirmed it disclosed sensitive customer data to an unauthorized third party after receiving fraudulent requests sent from a legitimate government agency email domain. The exposed data included birth dates, postal and email addresses, phone numbers, passports, driver’s licenses, verification selfies, account statements, and transaction histories. The domain looked right. The request came from a real government domain. The email gatekeeper failed anyway.
That is not a lesson. It is a cost. The lesson would require knowing exactly what validation step Revolut skipped, and we don’t.
Then there is Oracle. On September 14, 2026, reports surfaced that Oracle had begun another round of layoffs, again sending 6 a.m. termination emails to staff. Oracle’s workforce had already fallen by roughly 21,000 employees, or 13%, during fiscal 2026, and the company raised its estimated fiscal 2026 restructuring cost by $700 million, bringing the total to roughly $2.8 billion. I mention this because email is often the only channel companies use for high-stakes communication. If your email list is 40% disposable, your “important update” is not reaching humans. It is reaching inboxes that auto-delete in seven days.
I also spent time with a paper that hit HN the same week: Dream-RSI: Recursive Self-Improvement through Evolving Worlds (arXiv:2609.14858). The authors argue that fixed exploration strategies fail as search spaces scale, and that effective systems need to adapt their exploration based on accumulated history. Email validation has the same shape. A fixed regex and MX lookup was fine in 2010. In 2026, with burner domains, catch-alls, greylisting, and breached credentials, a fixed strategy fails. You need an API that treats validation as adaptive exploration: probe SMTP, check breach history, detect disposability, classify the provider, and synthesize a trust score.
The fourth source I read was harder to use. OpenReview’s paper on large language models developing novel social biases through adaptive exploration sat behind a browser verification wall, so I could not pull the full text. The title alone is enough to make me nervous: if adaptive exploration can introduce social biases in LLMs, adaptive validation can introduce false-positive biases in email scoring if we’re not careful. I’m still not sure if weighting is_free_email lower than a custom domain is the right call, or if it quietly discriminates against legitimate users on Gmail. That is an unresolved thought I’m leaving right here.
The composite score: 75 for test@gmail.com illustrates the tension. Gmail is a free email provider. The local part is test, a classic role-type string. MX is perfect. SMTP could not be verified in this call. Breach status could not be checked. So the API says: “This address is syntactically fine and the domain accepts mail, but I cannot confirm identity.” That is a 75. Not a 95. Not a binary valid.
If you are running lead scoring, that distinction is money. A 75 from a free provider with a role local is not the same lead as a 95 from a custom domain with SMTP proof and zero breaches. They should not get the same sales follow-up cadence or the same trial tier. They should not get the same trust.
Implications: what developers should actually do
So what do you ship?
Stop using MX found as a green light. Use it as a red-light filter only. If MX is missing, reject. If MX is present, keep investigating.
Build a trust ladder, not a trust cliff. The API gives you enough signals to tier users:
- Tier 1:
is_trusted_identity: true, custom domain, SMTP verified, no breaches, no greylisting. - Tier 2:
score70–90, free provider, SMTP unverified or catch-all unknown, no disposable flag. - Tier 3:
is_disposable: true,is_role: true, orscorebelow 70. Require extra proof.
Use syntax suggestions. The API returns a suggestion field for typos like gmial.com -> gmail.com. In my batch I saw a surprising number of near-miss domains. Fixing them at the point of signup recovers real users you would otherwise lose.
Watch catch-all and greylisting. A domain with is_catch_all: true will accept mail for any local part. That makes SMTP verification meaningless. A greylisted domain will defer your probe and make it look like failure. I wrote about greylisting separately in i validated 10,000 emails. the greylisting rate shocked me.. The short version: if you don’t handle greylisting, you discard valid addresses.
Use provider ID for B2B vs B2C segmentation. The email_provider field maps domains to Google Workspace, Microsoft, Proton, Zoho, Yandex, and others. A signup from a Google Workspace domain is not the same as a signup from a disposable Proton alias. Route them differently.
Treat breach status as a trust signal, not a punishment. An address with a high breach_count is not necessarily fake, but it is higher risk. Pair it with MFA prompts, password-strength checks, or step-up verification. The breach_status field is only useful if your HIBP key is valid, which brings me back to that error message: HIBP_API_KEY invalid or unauthorized. If you deploy this in production, configure the key. Otherwise you are flying blind on breaches.
Run composite checks before high-stakes actions. If you are about to send a layoff-style email, process a government request, or ship a password reset, do not rely on MX alone. Check is_trusted_identity. Check breach_count. Check is_role. The cost of a missed or spoofed address is too high.
I also keep thinking about a related problem: identity matching across sanctions and watchlists. In i ran 1,000 names through 2 ofac apis. 80 hits disagreed., I found that two reputable APIs disagreed on 8% of hits. Email validation has the same issue. No single signal is authoritative. The best you can do is stack signals and expose uncertainty.
The gap I’m leaving open
Here is the question I do not have a clean answer to: how should we weight a free-email address against a custom-domain address when both pass SMTP and neither is breached?
On one hand, a custom domain signals investment and identity. On the other hand, billions of real humans use Gmail. If your scoring model penalizes free email too heavily, you exclude legitimate users. If you ignore provider class entirely, you let burner-friendly providers blend in with business users.
The API gives you is_free_email and email_provider, but it does not tell you what to do with them. That is the gap. I’m still not sure if I should treat a verified Gmail inbox as equivalent to a verified custom-domain inbox, or if the custom domain deserves a trust bonus by default.
If you had a free weekend, what would you build with the Email Validator API’s is_trusted_identity composite and provider-ID signals? A tiered onboarding flow? A fraud-score microservice? A browser extension that colors email fields by trust level? I’d love to hear what you ship.
And if you want to run the same experiment, the API is the Email Validator API on RapidAPI, with code and issues on GitHub.
Top comments (0)