DEV Community

Tidy Tools
Tidy Tools

Posted on

Clean an email list without SMTP pings: typos, disposable and dead domains

Every outbound campaign starts with the same chore: a spreadsheet of addresses from a form, an event, a CRM export or a scraper, and the nagging feeling that a chunk of them will bounce.

Most "email verifier" tools answer this by connecting to the recipient's mail server and pretending to send a message (an SMTP "ping"). That used to work. Today it mostly does not: Gmail, Outlook, Yahoo and the big corporate gateways either block those probes or answer "yes" to every address, and the IP doing the probing can end up on a blocklist.

So I built a validator that deliberately skips SMTP and does everything that can be known without talking to a mail server. This post shows what that catches, what it cannot catch, and how to run it on a list or straight on a scraper's output.

What you can know without touching the mailbox

For every address, Email List Cleaner - Bulk Email Validator checks:

  • Syntax per RFC 5321/5322, including quoted local parts, internationalized domains and inputs like "Jane Doe <jane@example.com>" or mailto: links.
  • Does the domain exist, and does it accept mail? MX records over DNS-over-HTTPS, including domains that publish a null MX (an explicit "we take no mail") and the implicit fallback to A/AAAA.
  • Disposable domains: a community list of about 9,200 temporary-mail domains, plus disposable mail servers hidden behind custom domains.
  • Role, no-reply and free-provider addresses: info@, sales@, noreply@, Gmail, Outlook, GMX, QQ, Naver…
  • Typos: gmial.com, gmail.con, .cmo, with a didYouMean suggestion.
  • Mail provider from the MX hosts: Google, Microsoft, Proofpoint, Mimecast, Zoho and others.

Each address gets a verdict, a 0–100 score and a reason in plain English:

Verdict When What to do
valid Domain exists and accepts mail, nothing suspicious (role addresses stay valid with a lower score) Keep
risky Disposable, likely typo, no-reply, or several weak signals Fix the typo, or drop it for cold outreach
invalid Bad syntax, domain does not exist, null MX, no mail server Remove
unknown DNS did not answer even after a retry Run again later (not charged)

A real run

Input (you can also paste text, upload a CSV or point it at another run's dataset):

{
  "emails": ["jane.doe@gmail.com", "info@stripe.com", "someone@gmial.com", "test@mailinator.com"]
}
Enter fullscreen mode Exit fullscreen mode

A clean row from a CSV upload (1 October 2026), with the CSV's own Name and Company columns carried along:

{
  "input": "jane@apify.com",
  "verdict": "valid",
  "score": 90,
  "reason": "domain accepts e-mail; mailbox not verified (no SMTP check)",
  "domain": "apify.com",
  "hasMx": true,
  "mailProvider": "Google",
  "disposable": false,
  "role": false,
  "freeProvider": false,
  "mxRecords": [{ "priority": 1, "host": "aspmx.l.google.com" }, { "priority": 5, "host": "alt1.aspmx.l.google.com" }],
  "catchAll": null,
  "mailboxChecked": false,
  "charged": true,
  "csvRow": 2,
  "sourceFields": { "Name": "Jane", "Company": "Acme" }
}
Enter fullscreen mode Exit fullscreen mode

Note the honest bits: mailboxChecked: false and catchAll: null. Those can only be known over SMTP, so the tool says so instead of guessing.

The typo case is the one that saves real leads:

{
  "input": "someone@gmial.com",
  "verdict": "risky",
  "score": 20,
  "reason": "disposable (temporary) e-mail domain",
  "didYouMean": "someone@gmail.com",
  "reasons": [
    "disposable (temporary) e-mail domain",
    "possible typo: did you mean gmail.com?",
    "no MX record: mail would go to the domain's A/AAAA address (unusual for real mailboxes)"
  ]
}
Enter fullscreen mode Exit fullscreen mode

Other results from the same test:

  • user@example.com → invalid: "the domain does not accept e-mail (null MX record)"
  • zed@gmail.con → invalid: "the domain does not exist (NXDOMAIN)", with didYouMean: "zed@gmail.com"
  • noreply@github.com → risky: "no-reply or system mailbox"
  • bad..address@gmail.com → invalid: "consecutive dots in the local part", and not charged

Besides the table, every run saves two plain-text records, VALID_EMAILS and VALID_AND_RISKY_EMAILS, one address per line, ready to paste into your sending tool.

Validate a scraped lead list in one pipeline

Where this gets useful is right after a scraper. The Email Extractor / Website Contact Details Scraper returns rows with an emails array per website. Feed its dataset straight in; arrays are expanded and you can keep the website next to each verdict:

{ "datasetId": "<dataset ID of the contact scraper run>", "datasetField": "emails", "keepFields": ["url"] }
Enter fullscreen mode Exit fullscreen mode

Or from Python with the official apify-client package:

from apify_client import ApifyClient

client = ApifyClient("YOUR_APIFY_TOKEN")

contacts = client.actor("tidytools/website-contact-extractor").call(
    run_input={"domains": ["stumptowncoffee.com", "linear.app"]})

check = client.actor("tidytools/bulk-email-validator").call(
    run_input={"datasetId": contacts["defaultDatasetId"], "datasetField": "emails", "keepFields": ["url"]})

for row in client.dataset(check["defaultDatasetId"]).iterate_items():
    if row.get("success"):
        print(row["verdict"], row["email"], row.get("reason"))

clean = client.key_value_store(check["defaultKeyValueStoreId"]).get_record("VALID_EMAILS")
print(clean["value"])
Enter fullscreen mode Exit fullscreen mode

For a Google Sheet, publish it as CSV (File → Share → Publish to web → CSV) and paste the link into the CSV field. The column named like "email" is found automatically; set csvColumn if yours is called something else.

What it costs

$0.50 per 1,000 addresses on the Free and Starter plans ($0.45 on Scale, $0.40 on Business and higher). No start fee.

You pay only for unique addresses whose domain was actually checked. Free: syntax errors, unknown results, merged duplicates, and lines without an @ (empty CSV cells, notes). Every row says charged: true or false.

So a 10,000-address list costs at most $5. The run logs its worst case before it starts (Plan: 12 e-mail addresses × $0.0005 = at most $0.006), stops cleanly at your maximum charge per run, and if Apify restarts the run, addresses already done are not charged again.

What it cannot tell you

I would rather you know this before you rely on it:

  • valid does not mean "this mailbox exists". A made-up name at a real company domain (nobody123@stripe.com) comes back valid. If you need mailbox-level verification, use an SMTP verifier and accept its trade-offs.
  • Brand-new throwaway domains may not be on the disposable list yet.
  • Role addresses are not thrown out. info@ and sales@ often do receive mail; filter on role if you only want people.
  • Free providers are not "risky". freeProvider: true is just a flag, handy for separating B2B from consumer sign-ups.

Use the verdicts to remove what is certainly bad (invalid), review what is doubtful (risky), and keep the rest.

Responsible use

Only public DNS records are queried; no email is sent and no mail server is contacted, so the people behind the addresses are never touched. Addresses are sent to the validation backend only to be checked; they are not stored there and request contents are not logged. A valid verdict is not consent: validate only lists you are allowed to process, and follow GDPR, CAN-SPAM and CASL when you send.

Actor: https://apify.com/tidytools/bulk-email-validator


Disclosure: I built this Actor and the contact scraper, and I earn from their usage on Apify. All outputs above are real results from the Actor's test run of 1 October 2026; stripe.com, github.com and the other domains are public examples with no connection to me.

Top comments (0)