DEV Community

Himanshu mittal
Himanshu mittal

Posted on

How to verify lead domains and scrape metadata without paying SaaS credit markups (Python + n8n)

When building outbound lead pipelines or scaling agency data enrichment, domain decay is a silent budget killer: between 15% and 30% of scraped domains are dead, parked, or expired.

Running raw lists directly through credit-based enrichment tools burns paid credits on domains that do not even have active email servers[cite: 2].

To eliminate this waste, the core pre-flight enrichment pipeline can be handled with a zero-cost architecture using Python and a self-hosted n8n waterfall[cite: 1, 2].


1. Pre-Flight MX Validation via Google DoH (Zero Credits)

Direct SMTP handshakes on port 25 from cheap VPS hosts often get IP ranges blacklisted quickly.

Instead of paying per-row verification fees upfront, query Google’s public DNS-over-HTTPS API directly over standard port 443[cite: 2]:

https://dns.google/resolve?name=example.com&type=MX[cite: 3]

  • If the status is not 0 (NXDOMAIN) or has no mail servers, the domain is dropped instantly before wasting downstream resources[cite: 1, 3].
  • Dropping invalid records early avoids spending credits downstream while running entirely over routine HTTPS traffic[cite: 1, 2].

2. Lightweight Root Metadata Scraping (~200ms)

Spinning up headless browsers or paying for proxy-based scraping APIs creates heavy overhead just to extract company positioning.

For verified live domains, extract the homepage <title> and meta description in ~200ms using the pure Python standard library without browser automation overhead[cite: 1, 3]:

import re
import urllib.request

def fetch_metadata(domain):
    url = f"https://{domain}"
    req = urllib.request.Request(
        url, 
        headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64)'}
    )
    with urllib.request.urlopen(req, timeout=3) as response:
        html = response.read().decode('utf-8', errors='ignore')

    title = re.search(r'<title>(.*?)</title>', html, re.IGNORECASE)
    description = re.search(r'<meta\s+name=["\']description["\']\s+content=["\'](.*?)["\']', html, re.IGNORECASE)

    return {
        "title": title.group(1).strip() if title else None,
        "description": description.group(1).strip() if description else None
    }
Enter fullscreen mode Exit fullscreen mode

3. Structured LLM Opener Generation

Instead of passing massive HTML dumps into an LLM, feed only the clean metadata into gpt-4o-mini[cite: 1, 3].

Because the input payload is minimal, this drops enrichment costs from ~$180/1k leads down to just ~$2 in raw API tokens while generating contextual openers[cite: 1, 3].


Open Source Repository & Workflow

The standalone Python CLI tool and self-hosted n8n waterfall JSON workflow are open-sourced on GitHub[cite: 1, 3]:

πŸ‘‰ GitHub: https://github.com/mittalhimanshu76-vicky/coldengine-os[cite: 1, 3]

Top comments (0)