DEV Community

Cover image for How to Extract Public Emails and Phones From Websites Without a Paid API
Tim Zinin
Tim Zinin

Posted on Originally published at apify.com

How to Extract Public Emails and Phones From Websites Without a Paid API

The problem

Building a contact list from company websites sounds like a scraping task, but the naive version breaks in two predictable ways. First, volume: visiting a homepage, a contact page and an about page for each of 100 domains means 300 page loads, and copy-pasting addresses out of each one is not a job for a human. Second, correctness: the average company site mixes its own info@company.com with a webmaster's credit address, an agency's tracking script email, or a freelancer's portfolio link. Scrapers that lump every address they find into one column hand you a list that looks right and is quietly wrong.

Paid email-finder APIs take a different bet — they guess addresses from name patterns and charge per lookup whether or not the address was ever published. If what you need is what the company itself chose to publish, that is the wrong tool.

What the actor does

The Website Contact Scraper reads the server-delivered HTML of company sites — the same markup a search engine reads, with no browser and no JavaScript execution. Per domain it visits a bounded set of pages (homepage, /contact, /contact-us, /about, plus /impressum automatically for .de/.at/.ch sites; the list is configurable) and extracts:

  • Company-domain emails — addresses on the site's own domain, split into role-based and personal-style buckets by a deterministic heuristic;
  • Phone numbers — from tel: links, page text, or Organization JSON-LD, after placeholder filtering (posted, not carrier-verified);
  • Social profiles — Facebook, Instagram, LinkedIn, Twitter/X, YouTube, TikTok and WhatsApp links, accepted only on exact platform hosts.

Two design decisions do the heavy lifting for data quality. Third-party addresses found in the markup — an agency's script, a webmaster credit — land in a separate, clearly labelled thirdPartyEmails bucket, never in the company's own contact list. And every row carries page-level evidence: sourcePages and sourcePageUrls show exactly which pages contributed a contact, unreachablePages names the ones that failed and why, and partial: true plus an explicit error string tell you when contacts were confirmed on readable pages but part of the site could not be checked.

robots.txt is respected on every domain: disallowed target pages are never fetched (including cross-host redirect destinations) and are reported in robotsDisallowed. If robots.txt itself is inaccessible or rate-limited, the actor fails closed without fetching pages and returns a free partial row. Beyond the raw extraction fields, rows include decision metadata — a contactability score, evidence confidence with risks, recommendedAction, actionPriority, and safeToAutomate, which is always false: a published contact is evidence to review, not authorization to outreach.

Input is a list of up to 100 domains (or full URLs — the path is ignored), with optional checkPages, maxPagesPerDomain (default 5, max 6), and maxConcurrency (default 10). An alternative datasetId input mode lets another actor's dataset feed the domain list for chained workflows.

Example: input and output

{
    "domains": ["smithlawfirm.com", "apify.com"],
    "checkPages": ["", "/contact", "/about"],
    "maxPagesPerDomain": 3,
    "maxConcurrency": 1
}
Enter fullscreen mode Exit fullscreen mode

A real output row from the README:

{
    "input": "smithlawfirm.com",
    "domain": "smithlawfirm.com",
    "found": true,
    "emails": ["neil@smithlawfirm.com"],
    "thirdPartyEmails": [],
    "phones": ["+13147254400", "(314) 725-4400"],
    "socials": {
        "facebook": "https://www.facebook.com/neilsmithlaw",
        "linkedin": "https://www.linkedin.com/in/smithlawfirm",
        "youtube": "https://www.youtube.com/channel/UCwnmF7KeV9tejnAxREAhsRg"
    },
    "sourcePages": ["/"],
    "unreachablePages": ["/contact-us: http 404", "/about: http 404"],
    "robotsDisallowed": [],
    "partial": true,
    "error": "partial website check: 2 target page(s) could not be read; contacts may exist on content that was not available",
    "structuredDataFound": true,
    "checkedAt": "2026-07-30T07:43:30.604Z"
}
Enter fullscreen mode Exit fullscreen mode

Note the honesty of that row: found: true because a company-owned contact was confirmed, partial: true because two target pages 404'd — both facts travel together.

Pricing and the free limit

Pay per event: $0.005 per run start plus $0.002 per confirmed-contact row. A row is billed only when at least one company-owned contact was confirmed; every honest "couldn't check this one" row — JavaScript-only site, bot protection, robots refusal, non-existent domain — comes back free with the reason in error.

Apify's free plan gives $5 of usage credits per month. At this tariff, $5 covers up to 2,497 confirmed rows in a single run (0.005 + 0.002 × 2,497 ≈ $5.00) — or about 24 full 100-domain runs at $0.205 each.

Try it

Paste a domain list, start the run, and review the evidence fields before exporting anything: Website Contact Scraper — Public Emails & Phones

For AI agents and MCP

The actor takes JSON in and returns structured JSON out, with recommendedAction, failureType, retryable, and safeToAutomate=false fields built for machine routing — an agent can triage rows into review queues without a human reading HTML. The README documents the actor as a tool on the official hosted Apify MCP server, callable from ChatGPT, Claude Code, Claude Desktop, Cursor, and other MCP clients, and even ships a bounded instruction template for agents: keep only company-owned values, never treat thirdPartyEmails as company contacts, and route partial or safeToAutomate=false results to manual review.

Top comments (0)