DEV Community

Ancu Corp
Ancu Corp

Posted on

How to turn messy e-commerce reviews into a ranked defect table (no paid API)

I kept watching merchants burn hours copy-pasting reviews from Amazon and TikTok Shop into spreadsheets, trying to figure out why a competitor's product page converts better than theirs. The reviews are right there, in public. The bottleneck is never access — it's structure.

This is a walkthrough of the architecture I use to turn a messy review wall into a clean defect table you can actually act on. No paid API, no $99/mo scraping SaaS.

The three problems nobody mentions

Most "scrape reviews" tutorials stop at getting the HTML. That's the easy 20%. The real work:

  1. Pagination that lies. Review widgets lazy-load, virtualize, and silently cap at N pages. If you just loop ?page=1..50, you'll get the same 20 reviews fifty times and think you're done.
  2. Signal buried in noise. 4–5 star reviews are marketing. The actionable data is in the 1–3 star complaints — and those are the ones you need to cluster, not just count.
  3. Export that survives Excel. UTF-8 BOM or your Indonesian/Portuguese/Thai review text turns into éçã the moment a merchant opens the CSV.

Step 1 — Defeat the virtualized list

The reliable pattern is to scroll the container until the count stops changing, not to page-hack the URL:

def exhaust(container, expected, cap=1200):
    seen = set()
    stable = 0
    while stable < 3 and len(seen) < cap:
        nodes = container.query_selector_all("[data-review-id]")
        before = len(seen)
        for n in nodes:
            rid = n.get_attribute("data-review-id")
            seen.add(rid)
        container.scroll_by(0, 900)
        stable = stable + 1 if len(seen) == before else 0
    return len(seen)
Enter fullscreen mode Exit fullscreen mode

Three consecutive no-growth scrolls = you've hit the real bottom. The stable counter is what stops you both from quitting early and from looping forever on a widget that keeps re-rendering the same nodes.

Step 2 — Turn complaints into clusters, not counts

A star histogram tells you nothing. What you want is what people are angry about. A cheap, dependency-free approach: keyword buckets over the 1–3 star text.

CLUSTERS = {
    "packaging_damage": ["crushed", "dented", "arrived broken", "box damaged"],
    "dead_on_arrival":  ["doesn't work", "not working", "defective", "dead"],
    "wrong_sizing":     ["too small", "too big", "size chart", "doesn't fit"],
    "counterfeit":      ["fake", "not original", "knockoff", "replica"],
    "courier_delay":    ["late", "never arrived", "shipping took", "stuck in transit"],
    "poor_service":     ["no reply", "rude", "refund refused", "ignored"],
}

def classify(text):
    t = text.lower()
    hits = [k for k, kws in CLUSTERS.items() if any(kw in t for kw in kws)]
    return hits or ["other"]
Enter fullscreen mode Exit fullscreen mode

You don't need a transformer for this. Six buckets catch ~80% of real e-commerce complaints, run offline, and — crucially — are explainable to the merchant who's paying for the insight.

Step 3 — Export that opens correctly everywhere

import csv

def safe_csv(rows, path):
    with open(path, "w", newline="", encoding="utf-8-sig") as f:
        w = csv.writer(f)
        w.writerows(rows)
Enter fullscreen mode Exit fullscreen mode

utf-8-sig writes the BOM. That single change is the difference between a deliverable a merchant can open and one they'll email you back about.

The gotcha that wasted my afternoon

Amazon's review section injects the DOM after the network goes idle. page.wait_for_load_state("networkidle") fires before the reviews exist. You need to wait for the selector, not the network:

page.wait_for_selector("[data-review-id]", timeout=15000)
Enter fullscreen mode Exit fullscreen mode

If I had a dollar for every "empty result" bug that was really a race condition...

Where this gets you

Run this on 3–5 competitor products and you get a defect table like:

Cluster Competitor A Competitor B Your product
Packaging damage 41 12 ?
Dead on arrival 28 5 ?
Wrong sizing 9 33 ?

Suddenly "improve quality" becomes "fix packaging for the A-segment." That's the whole game — turning unstructured complaints into a ranked, specific action list.


I packaged the full dual-engine version of this (a Manifest V3 Chrome extension and a headless Python CLI, with the 6-cluster classifier and multi-sheet XLSX export built in) if you'd rather not rebuild it: OmniScraper AI — headless lead & review intelligence CLI.

But honestly — the three steps above will get you 80% of the way for free. Start there.

What's the messiest review source you've had to deal with? I'm curious whether TikTok Shop or Shopee is worse for hidden pagination.

Top comments (1)