I built a domain expiry monitor, and to test it against something other than my own domains I pointed it at the "our work" pages of 30 US web design agencies — the client sites they link to publicly — and checked every domain against its registry.
The results were mildly alarming, but the interesting part was everything that went wrong on the way there.
The results first
65 client domains across 14 agencies:
- 1 had already expired, 60 days earlier
- 3 were inside 45 days, the nearest with 3 days left
- the rest ran from 46 days to several years out, clustering at 3–6 months
I'm not naming the agencies or the domains. Telling someone privately that their client's domain is about to lapse is useful; publishing it is exposing somebody's client.
Why RDAP and not WHOIS
WHOIS returns free text in roughly as many formats as there are registrars. RDAP (RFC 7482/9083) returns JSON, over HTTPS, with a documented shape.
The single nicest property: a 404 is an answer. In RDAP it means "this domain is not registered", which is precisely the event worth alerting on. No string matching on "No match for domain".
const res = await fetch(`${base}domain/${name}`, {
headers: { Accept: 'application/rdap+json' },
})
if (res.status === 404) return { registered: false }
Three things that will bite you
1. The registrar's name is hidden in a jCard
There's no registrar field. The registrar is an entity with the registrar role, and its human-readable name lives inside a vCard-in-JSON structure:
const entity = (body.entities || []).find((e) => (e.roles || []).includes('registrar'))
const card = entity?.vcardArray?.[1] || []
const fn = card.find((f) => Array.isArray(f) && f[0] === 'fn')
const name = fn?.[3] ?? null // "GoDaddy.com, LLC"
vcardArray[1] is an array of ["fn", {}, "text", "GoDaddy.com, LLC"] triples. That awkwardness is why so many tools show you IANA registrar ID 146 instead of a name — and "where do I renew this" is the actual operational question for anyone holding domains across several registrars.
2. IANA's bootstrap does not cover every TLD
The bootstrap file maps TLDs to RDAP servers, and .de, .co and .gg simply aren't in it. There's no authoritative endpoint to ask.
The right response is to say so. "The registry publishes no expiry date" is a fact about the registry, not about the domain, and it's more useful to a user than a number you inferred from somewhere else. Any tool that confidently shows an expiry date for a .de is guessing.
3. Public suffixes
cartwright.co.uk naively reduces to co.uk. Query that and you get the registry's own record, then cheerfully report the registry's expiry date as your user's domain expiry. The full Public Suffix List is thousands of entries; even a small hardcoded set of co.uk, com.au, co.nz and friends prevents the embarrassing version of this bug.
The bug that nearly discredited the whole thing
My first extractor took every href on the page. For one agency it confidently reported licdn.com, hotjar.com, addtoany.com and gmpg.org as their clients.
All four came from the <head>: <link rel="preconnect"> hints and the XFN profile link WordPress writes into every page.
I started adding them to a blocklist, then stopped — every new analytics vendor adds another entry, and the blocklist can never be finished. The structural rule can:
// Only anchors. A person clicks anchors, and only anchors point at clients.
const re = /<a\b[^>]*?\bhref\s*=\s*("([^"]*)"|'([^']*)'|([^\s">]+))/gi
One line, and the entire category of failure disappears.
A related one: an agency's second domain isn't the agency's client. One firm's site was at atendesigngroup.com and they also owned aten.io; another at openmediafoundation.org owned open.media. Matching only the exact domain you crawled isn't enough — comparing the leading label catches both.
About a third of sites hard-block automated requests
This one changed the architecture. I wrote the fetcher with fetch, and got 403 on every single site, including with a complete set of genuine browser headers. Not a Cloudflare JS challenge you can wait out — a flat 403, identical in headless Chrome and in a real headed one. It's blocking by IP reputation.
I switched to driving a real browser, which fixed it for about two thirds of sites. That turned out to matter for a second reason I hadn't considered: most agency portfolios render their client grid client-side, in Webflow or React. Even a successful fetch would have returned markup with no client links in it, and I'd have spent a day blaming my extractor.
The unglamorous finding
Only 14 of 30 agency sites linked three or more client domains at all. The rest show client work as screenshots with no anchor.
That's a lot of backlinks being left on the floor.
If you want to look at what I built with this, it's at dropperch.com — free tier monitors domains. But the RDAP notes above are the genuinely reusable part, and they're yours whether or not you ever look at it.
Top comments (4)
"A 404 is an answer" holds only once you are sure you asked the right registry, and .com is the TLD where you cannot tell. Asking Verisign's
.comendpoint for a genuinely unregistered.comgives 404 withcontent-type: application/rdap+jsonand a zero-byte body. Asking that same endpoint forwikipedia.org- a live domain, wrong registry - gives byte-identical output: 404, same content type, empty body. Nothing in the response separates "not registered" from "you routed this to the wrong server".PIR's
.orgendpoint does it properly, for contrast: 404 with a 2938-byte RFC 9083 error object,errorCode 404,title "Object not found". So the discriminator exists in the spec, it just is not universally sent, and the registry that skips it is the one holding most of your client domains.This chains straight into your point 2, which is what makes it worth the guard. A bootstrap gap is exactly the condition that sends a query somewhere plausible but wrong, and the misroute then surfaces as the one event you alert on. Cheap check: treat a zero-length 404 as unknown rather than as an answer, and require a parsed
errorCodebefore declaring a domain unregistered. I only tested Verisign and PIR here, so I do not know how many other registries send the bare version.Hi Vinh — thanks for taking the time to test this properly.
Reproduced, and you are right. Same three probes:
verisign .com <- unregistered .com 404, 0 bytes, no errorCode
verisign .com <- live wikipedia.ORG 404, 0 bytes, no errorCode
PIR .org <- unregistered .org 404, 2940 bytes, errorCode 404
The first two are byte-identical, as you said. Since "not registered" is the one event my alerting exists to send, the failure mode is telling somebody their live domain has dropped — which is the worst wrong answer this thing can give.
You mentioned you had only tested Verisign and PIR, so I ran the same probe across everything in my bootstrap map. Verisign appears to be the outlier, not the norm:
Registry TLDs 404 body
Verisign .com .net 0 bytes
PIR .org 2940 bytes, errorCode 404
Identity Digital .info .io .me .ai 3258–3260, errorCode 404
Neustar .biz 3373, errorCode 404
Google Registry .dev .app 2564, errorCode 404
CentralNic .xyz 2134, errorCode 404
Radix .site .online 327–329, errorCode 404
NIXI .in 330, errorCode 404
Nominet .uk 69, errorCode 404
Fifteen TLDs, one bare responder. Unfortunately it is the one holding most of the names anyone asks me to watch.
That is also why I did not take the guard literally. Requiring a parsed errorCode before declaring a domain unregistered would mean .com and .net can never be reported as dropped — and that is the majority of what the product watches, so the check would remove the feature rather than protect it.
What I shipped instead aims at the condition you identified as the cause rather than at the symptom. A misroute needs a wrong endpoint, and my endpoints come from two places:
IANA's bootstrap, keyed by exact TLD. For these a bare 404 is still accepted — being wrong here requires the bootstrap itself to be wrong.
A small hardcoded fallback map for ccTLDs that run RDAP with no bootstrap entry (.io, .me, .sh). These are my guesses, and a TLD missing from the bootstrap is precisely the condition you describe.
So a bare 404 from group 2 is now inconclusive rather than an answer, and the verdict records which kind of evidence produced it (rfc9083 vs bare-404) so the alerting layer can be more cautious about the weaker one.
Residual exposure I have not closed: a bare 404 from Verisign via the correct bootstrap endpoint is still taken at face value. I do not have a way to distinguish it, and the endpoint is authoritative for .com, so the remaining risk is the bootstrap being wrong rather than my routing being wrong. If you know of a discriminator Verisign does send — a header, a differing response to a malformed name, anything — I would take it.
Thanks for this. It is a better catch than anything in the post.
No discriminator at the domain endpoint, as far as I can measure. Verisign is bare across the board: an unregistered
.com, a livewikipedia.org,notadomainwith no TLD at all, and-bad-.comall come back as 404 plusContent-Type: application/rdap+jsonplus zero bytes, and theirnameserver/andentity/misses do the same. The one shape boundary I did find is that requests which never reach the RDAP application answerHTTP/1.0 400 Bad requestwith no rdap content type at all — theirdomains?name=search and any unknown path segment behave that way — so you can separate a request the responder never saw from one it answered and refused, which is not the split you need.The part that may move your residual-risk line: Verisign serves
.comand.netfrom the same host under different path prefixes, and both are in the bootstrap.rdap.verisign.com/net/v1/domain/google.comis a live domain, the right vendor, and a bootstrap-listed endpoint, and it returns the identical bare 404. So the bootstrap can be entirely correct while the verdict is still wrong, if the TLD key your code derives from the name is off by one label. That makes one guard cheap: at request time, assert that the queried TLD matches the path prefix of the endpoint you are about to call, instead of treating a bootstrap hit as proof that the route is right.Some comments may only be visible to logged-in visitors. Sign in to view all comments.