DEV Community

Muhaymin Bin Mehmood
Muhaymin Bin Mehmood

Posted on

How I Built an Honest Website Image Auditor Without Pretending Estimates Are Measurements

Most image-audit tools end with one big, confident number: "You could save 73%!"

When I built BatchSet's Website Image Auditor, I kept asking what that number actually was. Was it measured? Estimated? Guessed from a file extension? Most tools don't say. The bytes they measured, the bytes they assumed, and the bytes they couldn't see all end up in one percentage.

My goal became simple: clearly separate what was measured, what was estimated, and what could not be measured. Validation later showed me where the first version still fell short. This post covers both the design and the part where my own numbers failed their test.

What the auditor does, briefly

You paste the URL of one public page. The server fetches that page's HTML, extracts the image references it can find, and tries to measure each image's transfer size. From that it produces:

  • findings per image (very large transfer, possible modern-format opportunity, missing width/height, missing alt),
  • a rules-based score out of 100,
  • the measured image bytes, with no projected saving,
  • and an option to re-encode up to 10 flagged images to WebP in your browser, download them as a ZIP, and see the actual before/after bytes.

Each of those steps involved a decision about what we can honestly claim.

1. Measure closer to what modern browsers receive

The first bug I hit was subtle. I was fetching images with a generic Accept: */* header, and Shopify stores looked terrible: almost everything came back as JPEG.

Many of those stores weren't serving JPEG to modern browsers. Their CDNs negotiate the format from the Accept header. Chrome says it accepts AVIF and WebP, so it gets WebP. My fetcher said "anything", so it got the legacy fallback.

The fix was to ask the way a modern browser does:

accept: "image/avif,image/webp,image/apng,image/*,*/*;q=0.8",
Enter fullscreen mode Exit fullscreen mode

This doesn't reproduce every browser decision. Viewport, device pixel ratio, client hints, caching and JavaScript can all change which asset a real visitor gets. But it's far more representative than requesting the generic fallback. It also makes the audit less dramatic: already-optimized sites score higher, and fewer images get flagged. An auditor that asks like a 2010 browser can find big "savings" almost anywhere.

Responsive images raised a related question. One product photo might appear in srcset as ?width=240, ?width=832 and ?width=3840. Counting those as three images inflates the total. Picking the thumbnail understates it. The crawler groups same-path variants and keeps the largest width it finds as a conservative stand-in. That doesn't prove every visitor downloaded that file, because the browser's final choice depends on viewport, DPR and layout. It's a deliberate "don't understate" choice, not a measurement of any one visit.

2. HEAD requests, and why "unknown" is not zero

When a server returns a trustworthy Content-Length for a HEAD request, we can read the asset's transfer size without downloading its body:

const res = await safeFetch(asset.src, { method: "HEAD", accept, allowedContentType: /^image\//i })
asset.bytes = res.contentLength ?? null   // no header → unknown, NOT 0
Enter fullscreen mode Exit fullscreen mode

That lets one scan size dozens of images within a serverless time budget. It isn't magic, though. HEAD can be missing, blocked or wrong, and some CDNs answer HEAD differently from GET.

The important part is the ?? null.

When there's no usable length, the easy (and wrong) move is to treat the image as 0 bytes. Every unknown image then quietly lowers the page's weight and raises its score, so a site that blocks your fetcher looks faster.

In the live auditor, an unknown size stays null, and an image without a measured size is left out of the analyzed set entirely:

  • it's excluded from the score and from the measured byte total,
  • the report shows found N images, analyzed M, so the gap is visible.

That choice has a cost, and you should know about it. Because the whole image is left out, its other issues aren't reported either, even observable ones like a missing alt attribute. It keeps the score coherent (every ratio is over the same set of images), but it means an unmeasurable image can hide a real accessibility problem. That's a trade-off I intend to revisit.

Separately, for the large research crawl behind our e-commerce image study, the collector goes one step further. If HEAD returns no length, it makes a one-byte Range request (Range: bytes=0-0) and reads the total size from the Content-Range header (bytes 0-0/48213). If the server ignores Range and replies 200, it takes Content-Length from the headers and closes the connection before the body arrives. If that fails too, the size is recorded as unknown and reported as a coverage figure, never zero-filled. That fallback exists in the research collector only, not in the live auditor.

How accurate is header-based sizing? On a set of 51 real product images, I compared the header-based sizes with fully downloaded bytes. The median difference was 0%. On that sample at least, this step is a real measurement.

3. Static HTML vs. the rendered page

The auditor reads the served HTML. It doesn't run JavaScript. It parses the HTML into a tree (no script execution, no sub-fetches) and looks for:

  • <img src> and srcset
  • lazy-loading patterns such as data-src and data-srcset
  • <picture><source srcset>
  • inline style="background-image: url(…)"

That's fast, cheap and deterministic. It also has blind spots, and the report should say what they are:

  • Images injected by JavaScript after load aren't in the served HTML.
  • CSS background images in external stylesheets aren't parsed.
  • Bot challenges. Some sites answer an unknown client with a challenge page such as Cloudflare's "Just a moment…". A naive crawler finds zero images there and reports a perfect score.

The last one worried me most, because it produces a confident wrong answer. The crawler now detects challenge pages (status 403/429/503, or known challenge markers in the HTML) and refuses to score them.

For those blocked pages, the auditor falls back to a browser-rendered image list from Google's PageSpeed Insights API, for the submitted page only. That gives network data: which images loaded, and how many bytes were transferred. It doesn't give a DOM, so there are no alt attributes and no width/height attributes.

The report shows that:

  • the badge changes from Full audit to Partial audit,
  • 3 of 5 checks run, and the alt-text and dimension checks are marked unavailable, not passed.

"Unavailable" and "passed" look similar in a UI and mean opposite things. Keeping them separate was one of the most important calls in the project.

A normal static scan: Full audit badge, 5 of 5 checks available, found 21 images and analyzed 21, measured bytes and no projected saving
A normal static scan: every check ran, so the badge says "Full audit".

A bot-protected site: browser-rendered fallback banner, Partial audit badge, 3 of 5 checks with rendered dimensions and alt text marked Unavailable
A site that blocked the crawler. The score is 100, but two checks couldn't run, and the report says so instead of counting them as passed.

4. A server that fetches URLs strangers type in

A website auditor is, structurally, an SSRF (server-side request forgery) machine. Someone types a URL and my server fetches it. If that URL is http://169.254.169.254/ or a hostname that resolves to 10.0.0.5, a careless fetcher will read cloud metadata or internal services.

Every request, for pages and for images, goes through one guarded fetcher:

  1. URL checks: only http/https, no embedded credentials.
  2. Resolve the hostname and block it if any of the resolved addresses is private, loopback, link-local or reserved.
  3. Pin the connection to the validated IP. The socket connects straight to the address we checked, and the hostname is used only for the TLS SNI and the Host header. There's no second DNS lookup, so a DNS-rebinding attacker has nothing to flip.
  4. Re-validate every redirect hop, with at most 3 redirects.
  5. Hard limits: a streaming byte cap, a timeout, and a content-type allowlist (HTML for pages, image/* for images).
const pinnedIp = await resolveSafe(u.hostname)   // throws if ANY address is private/reserved
client.request({
  host: pinnedIp,                 // connect to the checked IP, no re-resolution
  servername: u.hostname,         // TLS SNI + certificate check against the real name
  headers: { Host: u.hostname, "User-Agent": "BatchSetAuditBot/1.0 (+https://batchset.com)" },
})
Enter fullscreen mode Exit fullscreen mode

On top of that, the free public scan is deliberately small: one page, at most 50 images, a Turnstile check, and per-IP rate limits. Shared report links are unlisted, served with noindex, and expire after 7 days.

5. Why this is not a Lighthouse score

People see "score out of 100" and assume it's Lighthouse. It isn't, and it shouldn't be compared with one.

Lighthouse runs a simulated page load and scores page performance: timings like LCP and TBT. This auditor scores image delivery signals with a fixed rule set. Each category deducts up to a cap, in proportion to the share of analyzed images that trip it (for the two HTML-attribute checks, the share of analyzed <img> elements):

Category Max deduction Rule (as implemented)
Very large transfer 30 transfer size over 1 MB
Medium-heavy transfer 20 over 200 KB and up to 1 MB. The two size bands never overlap, so an image counts in one or the other
Possible modern-format opportunity 25 served as JPEG/PNG. Uses the response Content-Type, falling back to the URL extension only when there's no content type
Missing dimensions 15 <img> without width/height HTML attributes (not decoded or rendered size)
Missing alt 10 <img> without an alt attribute

Score = 100 − the deductions. For example, if half the analyzed images are over 1 MB, that category costs 15 of its 30 points.

Two honest caveats about these rules:

  • The format rule is a recommendation, not a verdict. PNG is a perfectly valid choice for some graphics, such as crisp UI assets or artwork that needs lossless edges. The rule flags an opportunity to check, not a defect.
  • The medium-heavy band is a byte threshold, not a comparison with the size the image is displayed at. In the report it's labelled as an estimate for that reason.

Given the same discovered markup and the same measurable asset data, these fixed rules produce the same score. There's no simulated network run involved. Real-world inputs still change, though: CDN responses, blocked requests and edits to the page can all move the score between scans. It's also not a Core Web Vitals measurement. Fixing a flagged image can help LCP or CLS if that image is part of the problem, but this report doesn't measure that.

6. Being a polite crawler

The fetcher identifies itself honestly as BatchSetAuditBot/1.0 with a URL.

What the robots.txt check covers, stated precisely: only the signed-in multi-page dashboard crawl checks robots.txt, and only for page URLs. Each page URL is checked before it's requested. Not covered: the free public scan (one URL a person typed in), the image requests that follow any scan, and the PageSpeed Insights fallback (Google fetches the page, not us) used when a site blocks our scanner.

  • The public scan fetches exactly one URL that a person typed in, plus that page's images, much like a browser opening a link, so it does not check robots.txt. It never follows links.
  • The signed-in multi-page crawl checks robots.txt before every page it requests. It follows the standard (RFC 9309) rules: the BatchSetAuditBot group takes precedence over *, the longest matching rule wins, Allow wins a tie, and * wildcards and $ anchors are supported. A disallowed page is never requested; it's listed in the report as "blocked by robots.txt". A redirect onto another origin gets that origin's robots.txt checked before we follow it. If robots.txt couldn't be read (a network error, timeout, server error, rate limit, or an HTML page served instead of a robots file), the crawl skips that site rather than assume permission. A missing file (404/410) means no rules. If the start page itself is disallowed, the audit stops there, with no fallback scan of any kind. Pages are fetched one at a time, about 0.75–1.25 s apart per origin (a redirect target is spaced on its own origin's clock), and longer if the site sends a Retry-After. A Retry-After over 10 seconds ends the crawl instead of waiting. A crawl covers at most 20 pages, and stops starting new pages after about 30 seconds so it fits the server's time limit; the report marks it as cut short. You can cancel a crawl mid-way: it stops the current request and the queue, keeps the pages it finished, and shows "Cancelled by user" without a score. It doesn't apply Crawl-delay values.
  • The research crawler visits hundreds of stores that nobody asked me to scan, so it runs more slowly. Sites are processed one at a time, with a jittered 1.2–1.8 s delay between page fetches to the same host (longer if Crawl-delay asks for it) and low concurrency for image requests. Its robots.txt handling is older and simpler than the parser above: it merges rule groups, ignores wildcard patterns, and treats a robots.txt it couldn't read as "no rules" instead of skipping the site. So before the study is published, I'm re-checking its crawl decisions against the new parser and excluding any page it shouldn't have fetched.

Confidence in the research data is computed by rules, never picked by hand. Each page gets a tier (high / medium / low / excluded) based on whether it's really a product page, how many images were analyzed, and coverage. Coverage is split into two numbers, because a single number misled me once:

  • analysis coverage = images sampled ÷ images discovered
  • size coverage = images sized ÷ images sampled

In an early pilot, one combined "coverage" figure showed a page as "100% covered" because every sampled image had been sized, even though we'd sampled only a slice of the page. Splitting the number fixed that. A page that hits the image cap can now never rate higher than medium.

7. The day my own estimate failed its test

The first version of the auditor showed an estimated saving before you converted anything, labelled "estimate" in the UI. It was a heuristic: the expected WebP re-encode ratio for each format, plus a resize factor for images served far wider than 2048 px.

For our e-commerce image study, I wanted to know whether that heuristic deserved to be in a published report. So I tested it properly. I took real product images from real stores, converted each one to WebP in real Chromium through the same code path the product uses, and compared the predicted saving with the actual one on held-out images.

It failed by a material amount:

  • Header-based sizing held up (the 0% median difference above).
  • The saving estimate didn't. The median error on held-out images was about 32 percentage points.
  • It overestimated PNG badly. It predicted savings in the mid-50s, while the actual median was around 18%.
  • AVIF re-encoded to WebP often got bigger. The estimator assumed "no gain", but the real result was a loss.
  • Small images (under ~20 KB) frequently grew when re-encoded.

I tried a revised model. When I checked it against a fresh, untouched set of stores, there weren't enough eligible images to reach a verdict either way. So I stopped. I didn't keep tuning until the numbers looked good.

What changed as a result:

  • The study publishes only measured facts: formats served, image bytes per page, share of pages with images over 1 MB, missing dimensions, missing alt text. It has no "you could save X%" headline, because I couldn't show that number is reliable.
  • The product no longer shows a projected saving. The report shows measured image bytes and rules-based findings, and says plainly that it doesn't project a percentage before converting. A saving appears only after images are actually re-encoded, labelled Measured after optimization.
  • The old heuristic survives in one place, and it's never shown as a number. When a signed-in crawl flags more than 10 images, it's a rough tie-breaker for picking which 10 to optimize first, and the UI says the selection was "prioritized by a rough estimate".
  • The real number comes from doing the work. When you fix images, they're re-encoded in your browser and you see the actual before/after bytes. That includes images that came out larger: they're shown as larger and counted against the total, not hidden.

What it doesn't do

A methodology post should list its limits, so here they are:

  • It scans one page (free), not your whole site.
  • robots.txt is checked only by the signed-in multi-page crawl, and only for page URLs. The free scan, image requests and the PageSpeed fallback aren't gated by it.
  • It doesn't execute JavaScript in the static scan, so images injected by JS and external-CSS backgrounds can be missed.
  • Images whose size can't be measured are left out of the analysis, including their alt/dimension findings.
  • The size variant we measure is a conservative stand-in, not proof of what a specific visitor downloaded.
  • In fallback mode, alt-text and dimension checks are unavailable.
  • The medium-heavy finding is a byte threshold, not a comparison with the displayed size.
  • The score is not Lighthouse and not Core Web Vitals.
  • It doesn't predict savings. The only savings figure comes from actually converting, and some images (small files, AVIF sources) can come out larger.

The takeaway

If you're building any tool that reports numbers about someone else's website, three rules did most of the work for me:

  1. Keep unknown as unknown. A missing size is not zero, and a blocked page is not a perfect page.
  2. Keep "unavailable" apart from "passed".
  3. Test your estimates against real outcomes, act on what you find, and say so publicly.

If you want to see how this looks on a real page, you can run a free scan of one of yours here, with no signup: BatchSet Website Image Auditor. If a number in the report looks wrong, tell me in the comments, because that's exactly the kind of bug I want to hear about.

Top comments (0)