A rule we hold to in Nakodo: if something can be read off the brand's own website, the setup flow finds it and shows it as found, with a Change button. It never shows an empty box and a label.
Prices, the pricing page, the social links and the product description all come from a research pass over the site. The newest one is the logo, which now sits beside the campaign name in the app. Nobody uploaded it and nobody pasted a URL.
The icons a page declares
Any site built in the last decade declares its icons in the head. Here is a real one, and you can run this yourself:
$ curl -s https://nakodo.app/ | grep -oE '<link[^>]*rel="[^"]*icon[^"]*"[^>]*>'
<link rel="icon" href="/favicon.ico" sizes="16x16 32x32 48x48"/>
<link rel="icon" href="/icon.svg" type="image/svg+xml"/>
<link rel="apple-touch-icon" href="/apple-touch-icon.png"/>
Three candidates, and the interesting question is which one you want. We draw the thing at 28 to 32 CSS pixels, next to a name, on screens that are mostly 2x. So:
$ for u in icon.svg apple-touch-icon.png favicon.ico; do
curl -sI https://nakodo.app/$u | grep -iE '^content-(type|length)' | tr -d '\r' | paste -sd' ' -
done
content-type: image/svg+xml content-length: 626
content-type: image/png content-length: 3785
content-type: image/vnd.microsoft.icon content-length: 2155
The SVG is sharp at any size and the smallest file of the three. The 180px PNG is sharp and six times the bytes. The ICO tops out at 48px, which is blurry on a retina screen at 32 CSS pixels.
So the wanted order is: the smallest icon that is still sharp, then larger sharp ones, then the largest of the blurry ones.
One comparator for two different preferences
That is two sort directions in one list, which is the only clever line in the file:
const SHARP = 64;
const rank = (size: number) => (size >= SHARP ? size : 10_000 - size);
Anything 64px or over sorts ascending by size, so the smallest sufficient icon wins. Anything under 64 sorts into the 9,900s, descending by size, so all of the blurry candidates land after all of the sharp ones and the best of a bad bunch comes first. No comparator with a branch, no two-pass partition.
Sizes come from the markup, with defaults where the markup is silent:
function iconSize(link: HTMLElement, rel: string[], path: string): number {
if (/svg/i.test(link.getAttribute("type") ?? "") || /\.svg$/i.test(path)) return SHARP;
const sizes = (link.getAttribute("sizes") ?? "").toLowerCase();
if (sizes === "any") return SHARP;
const listed = [...sizes.matchAll(/(\d+)x\d+/g)].map((m) => Number(m[1]));
if (listed.length > 0) return Math.max(...listed);
return rel.some((r) => r.startsWith("apple-touch-icon")) ? 180 : 16;
}
An SVG counts as exactly sharp rather than enormous, so it beats a 180px PNG without beating a 96px one for no reason. sizes="any" means the same thing. A link with no sizes is assumed to be a 16px favicon, except for an apple-touch-icon, which the convention says is 180.
Run the real function against the HTML above and it returns, in order: icon.svg, apple-touch-icon.png, favicon.ico. The declared favicon.ico with its sizes="16x16 32x32 48x48" ranks last, and the /favicon.ico that gets appended as a last-resort guess deduplicates against it through a Set.
Two kinds of icon we skip
for (const link of root.querySelectorAll("link[rel][href]")) {
const rel = link.getAttribute("rel")!.toLowerCase().split(/\s+/);
if (!rel.some((r) => r === "icon" || r.startsWith("apple-touch-icon"))) continue;
if (/dark/i.test(link.getAttribute("media") ?? "")) continue;
// ...
}
rel="mask-icon" never matches the allowed list, so Safari pinned-tab masks are out. Those are single-colour silhouettes meant to be tinted by the browser, and dropped onto a card they look like a smudge.
media="(prefers-color-scheme: dark)" icons are out too. An icon declared only for dark mode is drawn to sit on something dark. Our card is paper-coloured, and a white-on-transparent logo on paper is an empty square.
Both of these are the same mistake in different clothes: an asset whose correctness depends on a context you are not providing.
Every candidate is also forced to https (url.protocol = "https:"), because the app is served over https and a mixed-content image does not load. Sites that still declare http://cdn.example.com/icon.png get fixed up rather than skipped, and if the https version does not exist, the next candidate is tried.
Verify by fetching, at most three times
A URL in a link tag is a claim. The claim is checked:
const MAX_TRIES = 3;
export async function findLogo(candidates: string[]): Promise<string | null> {
for (const url of candidates.slice(0, MAX_TRIES)) {
const res = await tryFetch(url);
if (res?.status === 200 && /^image\//i.test(res.contentType) && res.body && res.url.startsWith("https:")) {
return res.url;
}
}
return null;
}
Four conditions, each earning its place: a 200 (a soft-404 HTML page is a very common answer for /favicon.ico), a content type that starts image/, a non-empty body, and a final URL that is still https, because a redirect can downgrade the scheme after we upgraded it. res.url is what gets stored, so a redirect is followed once here and never again by a browser.
tryFetch is not fetch. It is our hardened client, which refuses private address space at DNS resolution time, re-checks every hop of a redirect, and caps the response. Every URL in this feature came out of somebody else's HTML, and the user-pasted variant of the same function goes through exactly the same path:
export async function logoAt(input: string): Promise<string | null> {
let url: URL;
try { url = new URL(input.trim()); } catch { return null; }
if (url.protocol !== "https:" && url.protocol !== "http:") return null;
url.protocol = "https:";
return findLogo([url.toString()]);
}
Three fetches is a budget, not a limit on correctness. The whole lookup is also a nice-to-have, so it is started early and awaited late:
const logo = withTimeout(findLogo(home.icons), LOGO_MS, () => new Error("Logo timed out")).catch(() => null);
// ...sitemap, key pages, text extraction...
return { pages, logo: await logo };
It overlaps the sitemap read and the key-page crawl, so in the normal case it costs no wall-clock time at all, and in the bad case it costs eight seconds of a job that was already running. A site with no logo is a site with no logo, not a failed setup.
The React part: an image that failed before hydration
We store the URL, not the bytes. That is deliberate (we are not a CDN for other people's logos, and not storing a copy is one less thing to keep in step with the brand's own site), and it has a consequence: the image can start 404ing at any time, and when it does, a broken-image glyph next to a campaign name looks like our bug.
export function CampaignLogo({ src, size }: { src: string | null | undefined; size: number }) {
const [failed, setFailed] = useState<string | null>(null);
if (!src || failed === src) return null;
return (
<img
src={src}
alt=""
width={size}
height={size}
referrerPolicy="no-referrer"
className="shrink-0 rounded-md object-contain"
onError={() => setFailed(src)}
// One that failed before the page came to life never fires onError.
ref={(img) => {
if (img?.complete) img.decode().catch(() => setFailed(src));
}}
/>
);
}
Four decisions in twenty lines.
It is a plain img, not the framework's optimised image component, because that component only fetches from hosts you list in config and these hosts are every customer's domain. An allowlist you cannot write is not an allowlist.
referrerPolicy="no-referrer" means the brand's web server does not get a log line telling it which page of our app is showing their logo.
alt="" and return null on failure, because the logo carries no information the adjacent name does not. When it is missing, the name stands alone and nothing shifts.
And the ref is the one that is easy to get wrong. onError is a DOM event, so React only catches it if the element exists when the error happens. On a server-rendered page whose image is already cached as broken, the browser finishes and fails the request before hydration, the event has been and gone, and your onError handler never runs. img.complete is true for both loaded and failed images, so asking decode() and catching the rejection is the way to find out which. One line, and it closes the gap between "this component handles missing logos" and "this component handles missing logos on a reload".
What it looks like in practice
Campaign setup asks for a domain. The research pass comes back with a brief, the social links, the prices and now a logo, and the review step shows all of it as found, each with a Change next to it. The ⋯ menu on a campaign has a "Name and logo" item for the cases our ranking gets wrong, which takes a pasted URL and validates it with the same logoAt above, rate limited like every other write.
It is a small feature. It is also the difference between a setup flow that interrogates you and one that shows you what it already knows: the found state is the default, and the form is the exception.
If you want to see the found-not-asked idea from the outside, the how it works page walks through what the research pass reads off a site, and the pricing page is itself one of the pages that pass looks for on a customer's domain.
Top comments (0)