I built a personal shopping agent. It runs on my homelab, gathers sale items from a list of sources once a day with two refreshes later on, and a half-hourly job checks alerts and my watchlist. The point is to tell me when something I'd actually wear is reduced in my size.
The interesting part was never the agent. It was finding out how much of the retail web will talk to a program at all. This is the measured version: 85 sources, what each gave up, and four mistakes I made reading the results.
The scoreboard
Every number here comes from one run: the daily gather on 19 September, which took 854 seconds.
| Status | Sources |
|---|---|
| OK | 56 |
| Blocked | 22 |
| Parse failed | 4 |
| Reachable, nothing parsed | 1 |
| Error | 2 |
Only 45 of the 85 returned even a single item. Eleven of the "OK" sources answered fine and had nothing on sale that matched.
The run collected 4,112 discounted products. After dropping what wasn't menswear, wasn't clothing or footwear, wasn't in stock in my size or wasn't really reduced, 941 were in my size, plus 452 where the size couldn't be checked. The biggest single drop was "wrong size", at 1,411.
And the distribution is lopsided in a way I didn't expect. 21 Shopify stores produced 3,173 of the 4,112 items, more than three quarters, from one endpoint that has been sitting there the whole time.
Shopify's /products.json is the quiet hero
Any Shopify storefront will hand you /products.json: paginated, JSON, no key, no rendering, no scraping. Every product, every variant, with price and compare_at_price, so "is this actually reduced" is a subtraction rather than a guess about a strikethrough in the markup.
Critically, the variants carry per-size availability. That one field is the whole reason this project works.
The other pleasant surprise was a big retailer that ships a public, search-only Algolia key in its own sale page. Its site uses that key to power its sale filters. I use it the same way, from the same page, to ask for on-sale menswear in my sizes, and on that run it returned 400 items filtered server-side. No rendering, no parsing, no load a normal visitor wouldn't cause.
The pattern: before you write a scraper, check whether the shop's own front end is already calling something cleaner than the HTML.
The hard part isn't the deal. It's the size.
Eight sources list products and prices but have no per-size stock I can read: Nike, JD Sports, Selfridges, Footasylum, Puma, Converse, Clarks and one flash-sale site. Some don't put sizes on the listing at all. On others the size selector looks the same whether a size is in stock or gone. Puma loads its size grid and inventory from a later API call, so none of it is in the HTML.
So they return items I can't confirm I can buy. A 40%-off jacket that doesn't exist in my size isn't a deal, it's an errand.
The rule I settled on: keep them in the pool, tag them size_unknown, and never alert on them. They can turn up if I go looking, but they can never interrupt me. For any notifier, the real question is the false-positive rate, and what the user does after the third one.
Mistake one: I read a 403 as a "no"
Five well-known retailers were written off early as "robots.txt disallows". I'd built the polite thing — fetch robots, honour it, move on — and five big names said no, so I moved on.
They hadn't said no. Their robots.txt returns 403 to a plain HTTP request, because the file sits behind the same edge protection as everything else. My code caught the failure, couldn't parse any rules, and fell through to "assume disallowed". Fetched the way a browser fetches it, every one of those files allows *.
Failing to read a policy is not the same as the policy saying no. If your fallback is a decision, log it as a decision ("couldn't read robots, assuming disallow"), not as a fact.
The sting in the tail: I re-tested all five properly, and all five still gave me nothing. No men's prices on one, redirects and a region picker on two, an empty page on the other two. Right answer, wrong reason, for about a day. Those are the ones that rot quietly: the outcome looks correct, so nothing makes you check the logic.
Mistake two: "blocked" is three different things
When I probed a new batch of sources, I tried each one twice: once with a plain HTTP client and once rendered in a real browser. Same honest User-Agent, same robots check, no evasion. The only difference is that JavaScript runs.
| Shop | Plain HTTP client | Rendered in a browser |
|---|---|---|
| High-street fashion chain | connection timeout | 76 prices on the sale page |
| Off-price outlet | not tried | 287 prices |
| High-street chain | 200, but no product grid | 179 prices |
| Secondhand marketplace | 200 | 192 listings |
| Voucher aggregator | 200 | 230 KB of offers |
A plain client would have put the first and third in the blocked column. One was just rendering its grid in JavaScript, and the other didn't answer a bare client but was perfectly happy to serve a browser.
Two others, Sports Direct and a high-street chain, returned 404 to every URL I tried. That wasn't a block at all. I was guessing their category paths wrong, and Sports Direct parses fine now from its real sale page. That's the trap: in a health table, "edge-blocked", "renders empty without JavaScript" and "you asked for a page that doesn't exist" all show up as the same red cell. Only the first is the retailer's decision. The other two are my bugs, and they'll sit in the blocked column looking like someone else's fault.
So the status note is now worth more than the count. Record why a source is red, and re-test the reds with a different client now and then.
Mistake three: Foot Locker was my bug, not theirs
On the 19 September run, Foot Locker returned 48 discounted shoes and none in my size, because my reader couldn't find per-size stock. I'd filed it with the others: size selector, no availability markers, nothing to read.
The stock was there all along. Foot Locker ships exact per-size availability in JSON embedded in the product page, with UK, EU and US sizes for each variant. My reader only looked at the page's visible elements. I fixed it on 20 September to read the embedded JSON first and fall back to the old method, and Foot Locker went from 0 to 9 in-size deals on the next full refresh.
Same day, same lesson: some European stores list shoe sizes as bare numbers like 45.5, and my parser read them as UK sizes and threw them away as wrong. "They don't publish it" and "I didn't read it properly" look identical from the outside.
Mistake four: I wrote my own guess into the outfit builder
The daily digest includes a couple of outfit ideas. Early on it put a brown hoodie forward as something I'd wear. The model hadn't made that up. My own outfit-building code assumed brown and cream were colours I wear, because I'd written that assumption in without asking. That's a question for me to answer, not a default for the code to guess.
The builder now splits the work between code and model:
- Code picks the candidate pieces out of the pool, per slot (top, layer, bottom, footwear), inside my sizes. Sandals, mules and slides are excluded here, before the model sees anything.
- The model chooses ids from that list and writes the name and the one-line idea.
- Code validates the answer: the right slots filled, no more than one neutral, the cheaper look actually cheaper than the dearer one. If validation fails, it falls back to a deterministic pick.
The model does the part it's good at (taste, phrasing) and none of the part where being confidently wrong costs money.
Rules I gave myself
Worth stating, because the blocked column is 22 and the temptation to fix that is real:
- Identify honestly in the User-Agent. No pretending to be Chrome.
- No proxies, no captcha solving, no rotating anything.
- A challenge page, an error page or an empty render counts as blocked. Back off 72 hours, don't retry in a loop.
- Cache, and keep the number of requests per site per run small.
One retailer worked fine for about 30 page loads and then Akamai shut the door. That's the system telling you where the line is. Others sit behind DataDome and reCAPTCHA. They've decided, and the honest response is to record it and move on, not engineer around it.
Which leaves the real conclusion. The shops that were easiest to work with are the ones that already publish structured data, deliberately or as a side effect of their own front end. Everyone else is spending money on a wall, and the thing behind the wall — is this jacket in my size — is the only fact I ever wanted.
🤖 Drafted with AI assistance from my own homelab notes, logs and repos, then reviewed and edited before publishing.
Top comments (0)