"What is this site built with?" sounds like a one-off question, until you need the answer for 5,000 domains. Sales teams want every Shopify store in a list. Agencies want to audit prospects. Analysts want to count which CMS wins in a niche.
Here's how technology detection actually works, and what I learned building an open-source detector.
Fingerprints: the core idea
Every tool leaves traces. The open-source webappanalyzer project (a community continuation of the old Wappalyzer data) catalogs 7,600+ technologies with patterns like:
"WordPress": {
"meta": { "generator": "^WordPress ?([\\d.]+)?\\;version:\\1" },
"html": "<link rel=[\"']stylesheet[\"'] [^>]+/wp-(?:content|includes)/",
"implies": ["PHP", "MySQL"]
}
A detector loads the page and tests each pattern against:
| Signal | Example |
|---|---|
| HTML |
/wp-content/ paths |
| Script URLs | cdn.shopify.com/s/… |
| Meta tags | <meta name="generator" content="WordPress 6.8"> |
| Response headers |
server: cloudflare, x-powered-by: Next.js
|
| Cookies |
_shopify_y, hubspotutk
|
| DNS | MX records → Google Workspace or Microsoft 365, SPF includes |
| SSL issuer | Let's Encrypt, DigiCert… |
Then it resolves implies (WordPress → PHP + MySQL), requires and excludes, and keeps only matches above a confidence threshold.
The catch: half the fingerprints need JavaScript
When I counted, 3,342 of 7,628 technologies have js patterns: things like window.ABTasty, window.Intercom or AFRAME.version. These only exist after the page's JavaScript runs. A plain HTTP request never sees them, and neither do tools that load later through tag managers.
So I ran the same sites both ways:
| Site | HTML + headers + DNS | Headless Chrome |
|---|---|---|
| notion.so | 18 | 38 |
| allbirds.com | 16 | 29 |
| hubspot.com | 33 | 56 |
Roughly 2× more detections with a real browser: analytics, chat widgets, A/B testing, consent tools and front-end frameworks.
Tips if you build this yourself:
- Block images, fonts and media in the browser. You don't need them, and it makes each page much cheaper.
- Evaluate all JS chains in one
page.evaluatecall, wrapped in try/catch (some getters throw). - Also feed the browser's network requests into the script-URL and XHR patterns. Many tools load from tag managers you won't see in the initial HTML.
- Respect
robots.txt, and expect some big sites to answer 403 to automation.
Open source + a hosted version
The detector is open source (GPL-3.0, like the fingerprint data):
github.com/swiftkit-dev/website-tech-stack-detector
If you'd rather not run browsers yourself, the same code runs on Apify as
Website Tech Stack Detector: paste a list of domains,
choose fast or deep scan, and get JSON/CSV with technologies grouped by category, versions and the
email provider. It's pay-per-site, and unreachable sites are free.
What would you use bulk tech detection for? I'm curious which categories people care about most.
Written with AI assistance; the endpoints, numbers and examples were tested before publishing.
Top comments (0)