Every apparel store has a version of the same quiet problem: a product page
that is still up, still indexed, still taking traffic, and missing the size
half the buyers wanted. The product is not sold out. It is partly sold out,
which is worse, because nothing in the admin flags it.
I wanted a number instead of an anecdote, so I read the public catalogue of
every Shopify apparel store I could reach and counted.
The method, which is boring on purpose
Every Shopify store serves /products.json. It is public, paginated, and it
gives you variants with their options and their availability. No API key, no
app install, no scraping of rendered HTML:
import urllib.request, json
def catalogue(host, limit=250):
page = 1
while True:
url = f"https://{host}/products.json?limit={limit}&page={page}"
with urllib.request.urlopen(url) as r:
products = json.load(r)["products"]
if not products:
return
yield from products
page += 1
The hard part is not fetching. It is deciding what counts as a size run.
A variant's options are free text set by whoever built the store, so Size can
be size, SIZE, Talla, or option 2 instead of option 1, and the values can
be M, Medium, md, 8, 28x32, or One Size. A product is a size run
only if, after normalising, its variants differ along exactly one axis and that
axis orders. One Size is not a run. A colour axis crossed with a size axis is
several runs, not one, and has to be split by colour before it means anything —
otherwise a product with 3 colours and 7 sizes looks like a 21-size garment
with holes all over it.
Approached 438 stores. 354 answered, holding 122,425 products between
them. Of those, 298 returned readable size runs — 64,268 of them. The
products total belongs to the 354; the 298 hold 111,774 of them.
The finding
A run is broken when a size is sold out while sizes on both sides of it are
still in stock. Not "nearly sold out" — a gap.
Across buyable styles the broken rate is 10.3%. About one style in ten.
But the number that changed how I think about it is the shape, not the rate:
72.6% of broken styles are missing exactly one size.
The damage is not spread evenly. It is a single hole, almost every time, and it
is usually near the middle of the run, because the middle is where the demand
is and the middle is what sells out first.
XXS XS S M L XL XXL
✓ ✗ ✗ ✗ ✓ ✓ ✓
That shape matters for a practical reason. A problem that is one missing size
in one product is a line on a purchase order. A problem that is diffuse across
a catalogue is something you can only feel bad about. Most merchants have the
first kind and experience it as the second, because nothing tells them which
line it is.
The number that cuts against me, which I am putting here and not in a footnote
If nearly three in four broken styles are missing exactly one size, then a lot
of what I am calling "broken" is a run with one size standing — and a run with
one size standing is arguably not broken at all. It is sold through. Calling
that a problem inflates my own case.
So here is the conservative version. Styles with two or more sizes stranded
— genuinely stock that cannot move while the rest of the run sells — are
2.8% of buyable styles. Note the denominator: that is of all buyable
styles, not of broken ones. It is a much smaller number than 10.3% and it is
the one I would defend if someone pushed.
An earlier, smaller pass put that figure slightly lower, and merging every scan
moved it up — the direction that flatters me, which is the direction a
correction is easiest to skip. The research page linked below carries both
passes side by side. I would rather publish the conservative number than the
flattering one, partly because someone will re-run this and find out.
How many stores, which is the question people actually ask
Style rates are the honest unit but nobody thinks in them. The question is
"does my store have this", so here it is per catalogue, both ways, because the
two ways are fifteen points apart and I would rather show you that than pick
one.
| loose | conservative | |
|---|---|---|
| stores with at least one broken run | 78.5% (234/298) | 63.8% (190/298) |
| of the 101 largest catalogues | 100% (101/101) | 94.1% |
Loose counts any style with a hole. Conservative counts only the ones with two
or more sizes still stranded beside the hole — the same stricter definition as
the 2.8% above, lifted to the store level.
The gap between those columns is not noise, and it has a shape. It is fifteen
points across all 298 catalogues and six points among the largest 101: the
two measures converge as catalogues get bigger. A small catalogue has few runs,
so its single qualifying style is more likely to be one honestly sold down to
its last size. A large catalogue breaks runs faster than it retires them.
Which means most of what separates the flattering number from the defensible
one is small stores whose only broken run was, in fairness, finished. I lead
with 63.8% now. I led with the looser figure for ten days, and the research
page carries that correction rather than quietly swapping it out. It carries a
second one underneath: my size grammar could not read sizes spelled out in
full, and the day after I fixed that I re-derived every figure from the same
cached catalogues and the whole set moved up by a few tenths. Both corrections
are printed beside the numbers they changed.
The per-store rows are published, one line per catalogue, including the column
that makes the conservative count reproducible — so you can recompute either
number rather than take mine:
the dataset.
Why the admin doesn't tell you either way
Shopify's low-stock reporting is per-variant. It will tell you that
Shirt / Medium is at zero. It will not tell you that Shirt / Medium being
at zero is different from Shirt / XXS being at zero, because it has no model
of a run — no notion that these seven variants are a sequence with a middle,
and that a hole in the middle costs you the sale while a hole at the end mostly
does not.
The whole insight is that ordering the variants correctly is the analysis.
Once the sizes are in order, "is there a gap" is a one-liner:
def broken(in_stock):
"""in_stock is a list of bools, in size order."""
first, last = in_stock.index(True), len(in_stock) - in_stock[::-1].index(True)
return not all(in_stock[first:last])
Everything else — normalising Medium to M, deciding that 28x32 is a
waist-and-inseam grid rather than a run, splitting colour out — is the work
that makes that one-liner mean something.
Run it on your own catalogue
I packaged the scan as a free check, because the argument is much less
interesting than your own number: paste a store domain and it reads the public
catalogue and tells you which runs have holes and which size is missing. No
install, no account, nothing to connect.
https://sizecurve.bananafest-destiny.com/check
The full method, the samples, and every figure above with its working — plus
the figures I have withdrawn — are at
the research page.
If you have scanned catalogues this way and got a different rate, I would
genuinely like to know. My normaliser is opinionated, and size parsing is where
I would expect two honest implementations to disagree most.
Top comments (0)