DEV Community

Gio Rich
Gio Rich

Posted on Fully Autonomous

Detect a website's tech stack in bulk with Python (a Wappalyzer-style lookup)

I built this actor; it's a paid tool on Apify with a free trial credit.

"What is this site built with?" is easy to answer for one site with a browser extension. It gets annoying when you have 500 domains in a spreadsheet and want to know which ones run Shopify, which use HubSpot, or who's still on an old WordPress install.

I'm an 18-year-old engineering student, and I built Website Tech Stack Detector for bulk lookups like that. This post shows how to call it from Python and Node, what you get back, and how I'd use it to pull Shopify stores out of a domain list.

How it works

Each site gets one HTTP request with browser-like headers, plus DNS lookups (MX, TXT, NS, SOA, CNAME, A, PTR) and one TLS handshake. Headers, cookies, meta tags, script URLs, inline scripts and the HTML are matched against the open Wappalyzer rule set (7,600+ technology fingerprints). DOM rules are checked with cheerio instead of a browser, so analysis takes about 0.3 seconds per page.

The part I care about most: every detected technology comes with a confidence score and the evidence that matched (a header, a script URL, a DNS record...). You can check any result yourself instead of trusting a black box.

Python example

pip install apify-client
Enter fullscreen mode Exit fullscreen mode
from apify_client import ApifyClient

client = ApifyClient("<YOUR_APIFY_TOKEN>")

run = client.actor("rel8ble/website-tech-stack-detector").call(run_input={
    "urls": ["allbirds.com", "minimalistbaker.com", "nextjs.org"],
    "minConfidence": 50,
})

for site in client.dataset(run["defaultDatasetId"]).iterate_items():
    print(site["domain"], site["cms"], site["ecommerce"], site["emailProvider"])
Enter fullscreen mode Exit fullscreen mode

Node example

import { ApifyClient } from "apify-client";

const client = new ApifyClient({ token: "<YOUR_APIFY_TOKEN>" });

const run = await client.actor("rel8ble/website-tech-stack-detector").call({
  urls: ["allbirds.com", "minimalistbaker.com"],
  includeEvidence: false, // smaller output
});

const { items } = await client.dataset(run.defaultDatasetId).listItems();
for (const s of items) console.log(s.domain, s.technologyNames.join(", "));
Enter fullscreen mode Exit fullscreen mode

What comes back

One item per website. A real row for minimalistbaker.com, trimmed (the technologies array had 26 entries; I kept one):

{
  "domain": "minimalistbaker.com",
  "statusCode": 200,
  "blocked": false,
  "technologyCount": 26,
  "cms": ["WordPress"],
  "ecommerce": ["WooCommerce"],
  "analytics": ["Google Analytics"],
  "tagManagers": ["Google Tag Manager"],
  "advertising": ["Amazon Advertising"],
  "cdn": ["Cloudflare"],
  "programmingLanguages": ["PHP"],
  "server": "cloudflare",
  "tlsIssuer": "Google Trust Services / WE1",
  "tlsExpires": "2026-12-04T21:16:49.000Z",
  "emailProvider": "Google Workspace",
  "dnsTechnologies": ["Google Workspace"],
  "technologies": [
    {
      "name": "Cloudflare",
      "confidence": 100,
      "categories": ["CDN"],
      "source": "website",
      "evidence": [
        { "type": "headers", "key": "server", "match": "cloudflare" },
        { "type": "dns", "key": "soa", "match": ".cloudflare.com" }
      ]
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

The headline columns (cms, ecommerce, analytics, cdn, paymentProcessors, liveChat and so on) are flat arrays, so the dataset drops straight into a spreadsheet. dnsTechnologies is kept separate on purpose: a TXT verification record for Atlassian means the company uses Atlassian, not that Atlassian runs on the website.

Use case: find the Shopify stores in a domain list

Say you build Shopify apps or run a Shopify agency and you have a CSV of 1,000 ecommerce domains. You want only the Shopify ones, plus what analytics and email tools they use, so your outreach is relevant.

import csv
from apify_client import ApifyClient

with open("domains.csv", encoding="utf-8") as f:
    domains = [row[0].strip() for row in csv.reader(f) if row]

client = ApifyClient("<YOUR_APIFY_TOKEN>")
run = client.actor("rel8ble/website-tech-stack-detector").call(run_input={
    "urls": domains,
    "minConfidence": 50,
    "includeEvidence": False,
})

with open("shopify-stores.csv", "w", newline="", encoding="utf-8") as out:
    w = csv.writer(out)
    w.writerow(["domain", "analytics", "marketingAutomation", "paymentProcessors", "emailSenders"])
    for s in client.dataset(run["defaultDatasetId"]).iterate_items():
        if "Shopify" in (s.get("ecommerce") or []):
            w.writerow([
                s["domain"],
                "; ".join(s.get("analytics") or []),
                "; ".join(s.get("marketingAutomation") or []),
                "; ".join(s.get("paymentProcessors") or []),
                "; ".join(s.get("emailSenders") or []),
            ])
Enter fullscreen mode Exit fullscreen mode

emailSenders comes from the SPF record, so it shows services the domain authorised to send mail (Klaviyo, Mailchimp, SendGrid...). That's often a better signal of the marketing stack than what's visible on the homepage.

In my test run of 46 mixed domains, all 9 Shopify stores and all 9 WordPress sites were identified correctly, and 45 of 46 were analysed in 25 seconds (the one failure was a domain that doesn't exist).

What it costs

$3.00 per 1,000 results, one result = one website analysed. Sites that fail to load are saved as an error row for your records but never charged.

  • The 1,000-domain job above ≈ $3.00
  • 10,000 domains = $30
  • Apify's free plan gives $5 of monthly credit, about 1,600 sites

Limits

  • Homepage only (or the exact URL you pass). A tool that only loads on checkout won't be seen unless you give that URL.
  • No JavaScript execution. The big one: analytics tags loaded through Google Tag Manager can be missed. The actor sees GTM, but not always what's inside the container. Tags pasted directly into the page are detected.
  • Bot walls. Some big sites (G2, Indeed, some banks) return a challenge page. You still get headers, DNS and TLS with blocked: true.
  • Hosting hints are hints. Behind Cloudflare, the origin host is hidden and you'll see the CDN.
  • Fingerprints are signatures, not certainty. Use minConfidence: 50 for the strictest list.

Bug reports go to the Issues tab on the actor page, which I read.

Website Tech Stack Detector on Apify


This article was drafted with AI and published by me.

Top comments (0)