Here is a jacket's last 60 days of sales, by size:
| Size | Units sold |
|---|---|
| S | 10 |
| M | 30 |
| L | 24 |
| XL | 12 |
If you plan the next order from this table, XL gets 15.8% of it. XL is the weakest size
after S, and the report says so plainly.
But XL sold out on day 30. It sold 12 units in 30 days, which is the same rate as L. For
the other 30 days nobody could buy it. The report counts those days as zero, and a zero
that means "could not buy" looks exactly like a zero that means "did not want".
Statisticians call this censored demand. You saw the sales, not the demand, and the two
come apart exactly when a size runs out. The size that sells fastest runs out first, so
the error lands on your best sizes. The next order under-buys the size that ran out, it
runs out sooner, and the next report makes it look weaker again.
The fix is about twenty lines
If a size is out of stock now, and its last sale was a while ago, count it at the rate
it sold while it had stock:
from datetime import date
MAX_FACTOR = 2.0 # never more than double what was observed
MIN_GAP = 7 # out of stock for at least a week
MIN_UNITS = 3 # too few sales to say what a rate was
def adjusted(sizes, window_start, today):
"""sizes: {name: (units_sold, last_sale_date, in_stock_now)} -> {name: units}"""
span = (today - window_start).days + 1
out = {}
for name, (units, last_sale, in_stock) in sizes.items():
out[name] = units
if in_stock or units < MIN_UNITS or (today - last_sale).days < MIN_GAP:
continue
selling = max(1, (last_sale - window_start).days + 1)
out[name] = round(units * min(MAX_FACTOR, span / selling), 1)
return out
def shares(units):
total = sum(units.values())
return {k: round(100 * v / total, 1) for k, v in units.items()}
start, today = date(2026, 8, 1), date(2026, 9, 29)
jacket = {
"S": (10, date(2026, 9, 28), True),
"M": (30, date(2026, 9, 29), True),
"L": (24, date(2026, 9, 27), True),
"XL": (12, date(2026, 8, 30), False), # sold out on day 30 of 60
}
observed = {k: v[0] for k, v in jacket.items()}
print("observed", shares(observed))
print("adjusted", shares(adjusted(jacket, start, today)))
observed {'S': 13.2, 'M': 39.5, 'L': 31.6, 'XL': 15.8}
adjusted {'S': 11.4, 'M': 34.1, 'L': 27.3, 'XL': 27.3}
XL goes from 15.8% to 27.3%, level with L, which is what the sales said while XL had
stock. On a 120-unit reorder, that is 33 XL instead of 19.
Why each guard is there
The three constants carry the judgement, so here is the reason for each.
-
MAX_FACTOR = 2. A size that sold 3 units on day 2 and then ran out would otherwise be multiplied by 30. One good weekend is not a rate. Doubling fixes the common case, a size that ran out halfway through, and it caps the damage when the in-stock stretch was too short to measure anything. -
MIN_GAP = 7days. "Out of stock and last sold yesterday" usually means it sold out yesterday, and nothing has been lost yet. It may also be a restock that has not been counted in yet. A week of silence while out of stock is a stockout. A day is noise. -
MIN_UNITS = 3. Below that, there is no rate to extend.
And one thing it deliberately does not do: it never touches the return rate. If XL
comes back 30% of the time, that is a fact about the units that shipped, and invented
units should not change it.
What it cannot see
It only knows a size is out of stock now. A size that sold out in August and was
restocked in September looks fully in stock, and its gap goes uncounted. Seeing that
needs stock levels day by day, not just today's, and this reads only today's. So this
is the conservative half of the fix. It never adds demand that was not there, and
it misses some that was.
It also assumes the in-stock rate would have held. For a seasonal style, the second
half of the window may have been slower anyway. The 2x cap is what keeps that
assumption cheap when it is wrong.
This is how Sizecurve now
counts a sold-out size: it builds each style's size curve from sales net of returns,
and says how many to reorder in each size. On the purchase order it marks the line
"sold out since" and the date, so you can see which numbers were stretched and by how
much. I build it. It is a Shopify app with a 14-day free trial.
If you would rather try the arithmetic on your own numbers first, the
size split calculator does it in the browser,
with nothing to install. And on which size tends to go first, counted across 298 public
catalogues instead of assumed: which size goes first.
Top comments (0)