DEV Community

Cover image for I built a browser extension that identifies fonts inside images. Here's what the recognizer actually does
Zerrin Arslan
Zerrin Arslan

Posted on Originally published at fontboxdl.com

I built a browser extension that identifies fonts inside images. Here's what the recognizer actually does

Every "what font is this?" thread I have ever read ends the same way: someone posts a screenshot, three people guess, nobody is sure. The browser tools that exist (WhatFont and friends) read the CSS of a page, which is great until the lettering is a JPEG, a logo, a poster photo or a screenshot. Then they have nothing to read.

I run a font site, I had a catalog of a few hundred thousand fonts sitting in a database, and I had a 16 GB consumer GPU at home. I figured a font identifier that looks at pixels would be a weekend. It was a summer. This is what actually worked, what embarrassed me, and the numbers at the end.

Version 1: pretrained features and a nearest-neighbour index

The first version was the obvious thing. Render every font in the catalog into a few sample images, push them through an ImageNet EfficientNet-B2, store the embeddings, and at query time embed the user's image and return the nearest neighbours.

My self-test said 93% top-1. I was thrilled for about a week.

Then I built a benchmark from 137 real photos with known fonts (signage, book covers, packaging) and ran it. Top-5 accuracy: 0 out of 137. Not low. Zero. A hand-cropped, binarised, perfectly clean Montserrat photo returned garbage.

The self-test was lying because it used the gallery's own renders as queries. Same words, same rendering style. That measures near-duplicate retrieval, not font identity. An ImageNet backbone embeds "what word is this" and "what texture is this" far more strongly than "what are the letterforms". If you test an identifier, test it with words it has never seen and with photos you did not render yourself.

Version 2: teach the embedder what a font is

The fix was to fine-tune the backbone for font identity with an ArcFace head, on synthetic data I could generate at any scale:

  • around 35,000 font families as classes, with family-level labels (all weights of a family map to one class), because a photo of Montserrat Bold must not be a different class from Montserrat Regular
  • random real words from a 6k-word list, so the model cannot memorise the sample text
  • weight sampling along the variable-font axis, plus artificial thickening and thinning
  • augmentation that looks like the queries people actually send: perspective, uneven lighting, low resolution, JPEG artefacts, textured backgrounds
  • a later round added 12 "effect" renders per family: outline, 3D extrusion, glow, texture fill, emboss, pixelation, occlusion, wide tracking, multi-line layouts

Each round warm-started from the previous checkpoint. A pilot on 3,900 fonts took 23 minutes and moved unseen-word top-1 from 1% to 61%. The full run, about 5 hours on the home GPU, landed at roughly 74% top-1 and 86% top-5 on unseen words.

The number I care about most is generalisation to fonts the model never saw during training: about 59% top-1, against 65% for fonts it did see. Six points. That means adding a font to the index is "render 12 images, embed, append", no retraining. The index today holds 262,000 fonts and about 2.5 million render embeddings, and it is a plain tensor on CPU, not a vector database. Scoring is max(family prototype, best individual render). Median query latency is around half a second on 8 CPU threads; the GPU is shared with other jobs, so inference lives on CPU.

Things that broke in embarrassing ways

Tofu. Some fonts in a 170k gallery have no Latin glyphs at all (Khmer, Tamil, Arabic-only families). Rendered with Latin sample words, they produce rows of identical .notdef boxes. Those boxes became magnets: anything blocky matched them at 99%. I ended up scanning every gallery entry for renders made of near-identical connected components and blocking 453 of them. The lesson underneath: look at your gallery images, not just your metrics.

Weight blindness. Before family-level labels, a bold photo of a font with a regular-only class simply did not match. Obvious in hindsight.

The smart crop that was not. I wrote an automatic text-region cropper. On a ground-truth set it improved exactly 0 of 18 queries, but it pushed scores into the 95-98 range and doubled latency. Tight crops make everything look confident. It shipped turned off, and "score went up" is no longer accepted as evidence for anything.

Calibrating on 12 samples. I set a quality gate using a dozen hand-picked images. On 1,152 real queries it rejected 27% of legitimate uploads. Thresholds get calibrated on the full log now, nothing else.

Dark backgrounds. White text on a dark poster embeds differently from black on white. Inverting the image when its mean brightness is below 110 was a ten-line change and gave +2 correct on the 137-image benchmark. Cheapest win in the whole project.

Training on marketplace mockups. I tried adding designer preview images from a font marketplace as training data. The model learned the layout template of each designer instead of the typeface, and matched mockups to mockups. Preview images are search data, not training data.

Honest confidence

Cosine similarity is not confidence. A clipart image with no text at all can score 96%. So the ranking you see comes from vote aggregation across image variants (full image plus OCR line crops), not raw score, and the UI shows a "not sure, closest look-alikes" state when the top candidates cluster together instead of printing a fake 99%.

It is currently too cautious. In the comparison below it put the correct font at #1 on 9 of 12 images and still said "not sure" on 7 of them. That calibration is the next job, and it will be done on the real query log, see above.

The 12-image test

I rendered 12 known Google Fonts the way they show up in the wild: clean, textured, rotated, perspective-warped, blurred, low contrast, small JPEG, dark background. Then I uploaded the same 12 files to every free identifier I could reach through a browser. Score is the rank of the true family in the top 5.

Identifier Top-1 Top-5 Note
FontBoxDL (mine) 9 / 12 12 / 12
Open-source identifier model (ResNet18, ~1k families) 7 / 12 9 / 12 strong on photos, lost the clean serif and marker
WhatFontIs 6 / 12 10 / 12 OCR letter step; collapses on small dark JPEGs
Font Squirrel Matcherator 5 / 6 5 / 6 blocked after ~15 uploads, only 6 tested
WhatTheFont 0 / 12 0 / 12 commercial-only catalog, so it structurally cannot name a free font; its paid look-alikes were reasonable

It is my test, on my images, so read it as "the approach works", not as a league table. The honest weak spot: script fonts. Lobster on a textured, rotated background came back at #4 behind three look-alikes.

The extension

The browser extension is the thin part. Manifest V3, no host permissions. Right-click an image, pick "Identify font in this image", and the extension sends only the image URL to my server, which fetches the image itself (with the usual SSRF guards) and returns the matches. Nothing is read from a page until you click. Because people kept asking, it also grew a hover inspector that reports the font that actually rendered an element (from the browser's font-face set, not the first name in the CSS stack) and a "fonts on this page" list.

Porting it to Firefox was mostly manifest work: background.scripts instead of a service worker, a browser_specific_settings.gecko block, and the data-collection declaration AMO now requires.

Links, if you want to poke at it: the web tool with cropping, the Chrome extension, the Firefox add-on page, and the longer, less technical write-up with five worked examples: How to identify a font from an image.

What kinds of images stump the identifiers you have tried? Hand-lettering, neon, chrome effects, tiny screenshots? Misses are the only training data that matters at this point, so I would genuinely like to see them.

Top comments (0)