Writing a regex that finds email addresses on a web page takes about a minute. Writing one whose output you can
paste into a CRM without apologising takes considerably longer, and the difference is entirely in what you
discard.
I built a contact extractor recently and ran it against real sites. Everything below is a false positive it
produced on the first pass, and what it took to stop each one.
heretohelp@stripe.comt
That extra t is not a typo in this post. The page had an address immediately followed by another word, and the
regex ran straight through the boundary.
The obvious guard is to check the domain has a valid public suffix, and the obvious way to do that is a library
like tldts. Except the naive call does not help:
import { getDomain } from 'tldts';
getDomain('stripe.comt'); // 'stripe.comt' — happily accepted
By default these libraries treat any trailing label as a suffix, because new top-level domains appear all the
time. What you want is the flag that says the suffix must be one ICANN actually publishes:
import { parse } from 'tldts';
parse('stripe.comt').isIcann; // false
parse('stripe.com').isIcann; // true
One property, and a whole class of junk disappears: addresses glued to the next word, filenames like
logo@2x.png, anything ending in an invented suffix.
+33 1 00 00 00 00
Four French phone numbers, all valid according to libphonenumber, all on the same page. They were on Stripe's
French sales page, which publishes dial patterns showing how regional numbers are formatted. libphonenumber
validates them because structurally they are perfectly good numbers. They are simply not anyone's phone.
The rule that catches them without catching real numbers:
const national = parsed.nationalNumber;
if (/^(\d)\1+$/.test(national) || /(\d)\1{5,}/.test(national)) return; // all one digit, or six in a row
Real numbers do not have six identical digits in a row. Placeholders almost always do.
+49123456789012 from the text "VAT 123456789012"
A twelve-digit run in prose, which libphonenumber was willing to read as a German number. Also seen: order
numbers, registration numbers, and anything else long and numeric.
The fix is not a better validator, it is a question about the source. A phone number written for a human to
read has punctuation in it. A VAT number does not:
// Candidates found in prose must look like a written phone number.
if (source === 'text' && !/^(\+|00)/.test(candidate) && !/[\s().-]/.test(candidate)) return;
Numbers taken from a tel: link skip this check, because there the site itself has declared what it is.
u003esales@stripe.com
This one only appeared when the scraper ran on a datacenter IP, because the site served different HTML there.
The u003e is a JSON-escaped > sitting in front of a real address inside a <script> tag.
The root cause is a detail of cheerio that is easy to miss:
$('body').text() // includes the contents of <script> and <style>
On a modern site that means you are scanning large JSON blobs, complete with escape sequences and configuration
data. Extracting from what a visitor can actually read fixes it:
const body = $('body').clone();
body.find('script, style, noscript, template, svg').remove();
const text = body.text();
That change also removed several phone numbers on another site that had come from embedded configuration rather
than from anything on the page. Fewer results, all of them real, which is the right trade for a lead list.
The same lesson, in a different format
PDFs have the mirror-image version of this problem. The library gives you positioned text fragments, not lines,
so the naive join produces one run-on paragraph:
content.items.map(i => i.str).join(' ') // goodbye, document structure
Fragments that share a baseline are one line, and a vertical gap noticeably larger than the page's usual line
spacing is a paragraph break. Reconstructing that takes maybe thirty lines of code and is the difference between
text a language model can use and text it will misread.
Two estimators inside that turned out to matter more than expected:
- Line spacing should be a low percentile of the gaps, not the median. On a page of prose both work. On a title page or an invoice most gaps are paragraph gaps, and the median lands on one of them.
- Body text size, used to decide which lines are headings, should be the most common size, not the median. Body text is by definition the size that repeats. A median is dragged upward by a page full of headings.
And a heading needs more than a large font. It is also short, does not end mid-clause, and does not start
lower-case. Font size alone promotes every wrapped sentence that happens to sit in a slightly larger face.
The general point
In extraction work, precision is the product. Anyone can raise recall by loosening a pattern; the value is in
the rules that decide what not to return, and every one of those rules should come from a real document that
broke something rather than from imagination.
Both tools are on Apify if you want the output rather than the code:
Website Contact Extractor and
Document Text Extractor. The source for both is on
GitHub, including the tests that pin each of the cases above.
Top comments (0)