I've been writing a checker that looks at the website a local business lists (salons, barbers, plumbers) and decides whether it works. The first version was simple: fetch the URL, and if that fails, call the site broken.
Then I ran it on about 50 salon, spa and barber websites in Nairobi and read every verdict by hand. A few "broken" sites worked fine in a browser, one site came out dead in one run and healthy in the next, and a parked domain looked healthy. Below are the cases and what the checker does now. Everything here is Node 20 with the built-in fetch.
The helper used in the snippets:
async function fetchPage(url, init = {}) {
try {
const res = await fetch(url, { redirect: 'follow', signal: AbortSignal.timeout(15_000), ...init });
return { ok: res.ok, status: res.status, finalUrl: res.url, body: await res.text() };
} catch (err) {
return { ok: false, status: 0, code: err.cause?.code ?? err.name };
}
}
1. HTTP says 404, HTTPS works
I request http:// first, because whether plain HTTP redirects to HTTPS is worth knowing. One barbershop's server answered plain HTTP with a 404 and no redirect, while https:// served the real site. Chrome has defaulted to HTTPS for typed addresses since version 90, and search results link to HTTPS, so almost nobody sees that 404. Calling the site broken was wrong.
Now an error over HTTP only counts if HTTPS fails too. The missing redirect is still recorded, as its own smaller finding:
const http = await fetchPage(`http://${host}${path}`);
const redirected = Boolean(http.finalUrl?.startsWith('https:'));
let page = http;
if (!http.ok && httpsWorks && !redirected) {
// Plain HTTP answers with an error page, but HTTPS visitors get the site.
page = await fetchPage(`https://${host}${path}`);
}
const httpRedirectsToHttps = http.status > 0 ? redirected : null;
httpsWorks comes from an earlier request (I fetch https://host/robots.txt first, which doubles as the HTTPS check). The !redirected guard matters: if HTTP already redirected to HTTPS and that page is a 404, fetching it again tells you nothing new.
2. The listed link is dead, the site isn't
Business listings often point at a deep page rather than the homepage. One salon's listing linked to /beauty-salon, which returns 404, while the homepage works. "Your website is broken" is the wrong verdict. "Your listing links to a missing page" is the right one, and it's a smaller fix for the owner.
Falling back to the homepage is only safe when the domain belongs to the business, though. Another listing pointed at a page on a booking platform. That platform's homepage works, and it says nothing about the salon. So the fallback only runs when the domain carries a distinctive word from the business name:
import { getDomain } from 'tldts';
const GENERIC = new Set(['salon', 'hair', 'beauty', 'barber', 'barbers', 'studio', 'spa', 'nails', 'house', 'shop']);
function domainCarriesName(name, url) {
const label = (getDomain(new URL(url).hostname) ?? '').split('.')[0];
const words = name.toLowerCase().split(/[^a-z0-9]+/).filter((w) => w.length >= 4 && !GENERIC.has(w));
return label.length >= 4 && words.some((w) => label.includes(w));
}
domainCarriesName('House of Treasures Beauty Salon', 'https://houseoftreasureskenya.com/beauty-salon'); // true
domainCarriesName('Posh Palace Hair Studio', 'https://booking-platform.example/p/12932'); // false
The generic-word list is what keeps "Glam Hair Salon" from matching a domain like salonbookings.com.
3. "Page Not Found" with a 200
That booking-platform link from case 2 returned status 200, with <title>Page Not Found</title> and almost no text. The status code says fine, the page says gone. On short pages I now check the title:
const NOT_FOUND_TITLE = /(^|[^a-z0-9])(404|page not found|not found|page (could not|cannot|can't) be found)([^a-z0-9]|$)/i;
function isSoft404($, visibleText) {
return visibleText.length < 3000 && NOT_FOUND_TITLE.test($('title').first().text().trim());
}
The length guard keeps a real page from being flagged because its title happens to contain "not found". The check also runs after the parked-domain and default-server-page checks, so those keep their more specific labels.
4. A certificate chain missing its intermediate
One server sent only its own certificate, without the intermediate that links it to a trusted root. openssl s_client shows it plainly: one certificate in the chain and Verify return code: 21 (unable to verify the first certificate).
Browsers usually cope. Chrome and Edge fetch the missing intermediate from the URL inside the certificate, and Firefox ships a list of known intermediates. Node's fetch on a Linux server doesn't, and fails with UNABLE_TO_VERIFY_LEAF_SIGNATURE. (The same request succeeded on my Windows laptop with Node 24, so test where you deploy.)
For that one error code, the checker reads the page anyway and records the certificate problem as a finding instead of "down":
import { Agent, fetch as undiciFetch } from 'undici';
const lenient = new Agent({ connect: { rejectUnauthorized: false } });
let page = await fetchPage(url);
if (page.code === 'UNABLE_TO_VERIFY_LEAF_SIGNATURE') {
const res = await undiciFetch(url, { dispatcher: lenient, signal: AbortSignal.timeout(15_000) });
page = { ok: res.ok, status: res.status, finalUrl: res.url, body: await res.text(), certIssue: 'incomplete chain' };
}
This only reads public HTML and sends nothing. It's deliberately narrow: an expired, self-signed or wrong-host certificate makes browsers show a full-page warning, so those stay "broken".
5. Slow is not dead
One WordPress site took 34, 17 and 11 seconds to answer three requests in a row with curl. With a 15-second timeout and one retry, it "loaded in 8.6 s" in one run and came out as "not responding" in the next. Same site, opposite verdicts.
Now a page that times out gets one more try with 45 seconds, and the load time becomes a finding of its own:
let page = await fetchPage(url);
if (page.code === 'TimeoutError') {
page = await fetchPage(url, { signal: AbortSignal.timeout(45_000) });
}
The first probe (robots.txt) keeps its short timeout and no long retry, so a host that doesn't answer at all still fails fast instead of holding up the batch for a minute.
And one the other way round
A parked domain returned 200 with a 114-byte page. There was no "this domain is for sale" text to match, only a script:
<script>window.location.href="/lander"</script>
Registrar parking pages often look like this. A tiny page whose only job is to send the visitor to /lander now counts as parked:
const PARKED_LANDER = /(window\.)?location(\.href)?\s*=\s*["']\/lander\b/i;
const isParked = html.length < 3000 && PARKED_LANDER.test(html);
Where it ended up
After these changes, 13 of the 57 businesses in that batch that list a website really did have a broken one: 10 domains no longer resolve, one is parked, one listing points at a booking page that's gone, and one server shows its default page. Four of the 13 have between 79 and 178 reviews each, so they're busy businesses still sending visitors to a dead link.
The general lesson for any uptime-style check: a failed request is a fact about your request, not yet a fact about the site. Check the other scheme, the homepage and the certificate, and give a slow site a second chance before calling it dead.
Top comments (0)