DEV Community

Cover image for We built a deal site that refuses to show fake discounts. Here's what broke along the way.
WhatNotSell
WhatNotSell

Posted on

We built a deal site that refuses to show fake discounts. Here's what broke along the way.

Most deal sites have the same problem: the discount is whatever the store says it is. "Was $299, now $149" gets published as 50% off, even when nobody ever paid $299.

We built WhatNotSell to do one thing differently. We only show a discount when the store's real original price backs it up. No estimates, no "compare at" guesses, no inflated percentages. If a product has no real higher original price, it shows no discount at all.

That rule sounds simple. Enforcing it across hundreds of thousands of product feed rows a day, from eight different sources, turned out to be the whole job. Here's how the system works and the mistakes that taught us the most.

The stack

  • Next.js 16 (App Router, Turbopack) on Vercel
  • Supabase for Postgres and auth
  • GitHub Actions plus Vercel Cron to run the imports
  • Product feeds from affiliate networks (Awin, CJ, Impact, Rakuten and others), plus eBay and Amazon

Today that adds up to roughly 11,000 live items from 180+ stores, refreshed every day.

Rule #1: one function decides every discount

Every import route, no matter which network it reads from, has to run its discount through the same function:

export function honestDiscount(price: number, originalPrice: number): number {
  if (!price || price <= 0) return 0
  if (!originalPrice || originalPrice <= price) return 0
  return Math.round(((originalPrice - price) / originalPrice) * 100)
}
Enter fullscreen mode Exit fullscreen mode

That's it. No real higher original price means 0%, and a 0% item is treated as a plain catalog listing, not a deal.

It sounds almost too basic to write about. But during a review we found that two of our own import routes weren't using it. They were calculating a discount some other way, which meant we were doing exactly the thing the site exists to prevent. Nothing crashed. The numbers just looked plausible.

Lesson: a rule that matters has to live in one function, and every new or changed code path has to be checked against it. "We have a helper for that" is not the same as "everything uses the helper."

Rule #2: we replaced an AI classifier with boring rules

Every product needs a category (Electronics, Home, Fashion and so on). Feeds are messy: some give you a category, some give you a merchant-specific taxonomy, some give you nothing but a title.

Our first version sent uncertain products to an LLM to classify. It worked, but two things bothered us:

  1. Cost grew with volume. More feeds meant more calls, every day, forever.
  2. Wrong answers were hard to fix. When it filed a drill under "Home", there was no rule to correct. We could only hope it did better next time.

Once we understood the patterns, we replaced it with a deterministic pipeline. First match wins:

  1. Merchant pin (a single-category store)
  2. Strong title keyword (unambiguous product names)
  3. The feed's own taxonomy
  4. Fetch context (the search the item came from)
  5. Weaker title keywords
  6. The merchant's default category
  7. Misc

All the keywords, pins and categories live in database tables, not in code. There's no cache, so fixing a rule corrects every affected product on the next import. Classification now costs nothing to run, and when something is wrong we can see exactly which layer made the call.

Lesson: an LLM is a great way to discover the rules. For a repetitive, high-volume job, it's often better as the prototype than as the permanent solution.

Rule #3: scores are recomputed, never nudged

Each deal gets a 0 to 100 score from the real discount, how much we trust the retailer, price history, and community votes. The one design decision we'd repeat: the score is a pure function, recomputed from scratch every time. Votes are stored separately and folded in on each recompute. Nothing ever does score += 5.

That makes the score reproducible and debuggable. If a number looks off, we can recompute it from its inputs and see why.

The failures that taught us the most

None of our worst bugs threw an error. They all failed silently.

An import that vanished for two months. Our Amazon import was accidentally dropped from the daily GitHub Actions schedule during an unrelated edit. Nothing failed, because nothing ran. Our monitoring checked for runs that went stale, but a job that's never started doesn't produce a stale run. We found it by noticing the Amazon section looked thin.

Runs that never finished. One feed request had no timeout. When the network hung, the import sat in a "running" state indefinitely. These piled up for about two months before we caught them. The fix was a timeout on every external fetch, plus an automatic close for any run stuck too long.

A green workflow hiding red endpoints. Our GitHub Actions job calls about 30 endpoints in sequence. If one returned a 504, the workflow used to finish green anyway, so the only way to catch a failure was to read the log. Now any non-2xx response fails the run, and GitHub emails us.

Two jobs fighting each other every day. This one we found just this week. Our import classifies each product using the feed's own category data. A separate nightly job re-checks every product's category using only its title, so rule fixes spread without waiting for the next import. The problem: the nightly job had less information than the import, so for about 280 products a day it replaced a correct category with a worse guess (laptops moved from "Computers & Laptops" to "Electronics", drills to "Home"). The next import fixed them, and the nightly job broke them again. Since the nightly job ran last, the wrong answer was what shipped.

The fix was to make the classifier report which layer produced its answer. The nightly job now only overrides a category when its answer comes from a layer that outranks the feed's own data. Daily changes dropped from about 280 to 17.

Lesson: for a data pipeline, "no errors" means very little. We now keep a run ledger for every import (what ran, how many rows, how many skipped and why) and alert on anything unusual: a source going quiet, a run stuck open, skip counts that jump. Most of our real bugs were found by reading those numbers, not by an exception.

What the data says

Because every discount on the site is backed by a real original price, the data is useful on its own. Across the 3,543 deals we had with verified discounts at the end of September, the average real discount was 36% and the median 33%. It varies a lot by category: laptops averaged just 17% off, electronics 27%, men's clothing 51% and fragrance 70%. Only 5% of genuine discounts reached 70% or more.

That's the practical takeaway for shoppers: a "70% off" banner is far more likely to be measured against an inflated price than to be a real 70% drop. We keep a live version of these numbers in our Deal Index.

If you're building something similar

  • Put your most important business rule in one function, and audit every code path against it.
  • Use AI to learn the rules, then consider replacing it with the rules.
  • Make derived values (scores, categories) pure functions of their inputs.
  • Assume your pipeline fails silently. Monitor for missing activity, not just errors.
  • When two jobs write the same field, make sure the one with less information can't overrule the one with more.

WhatNotSell is live at whatnotsell.com. If you want to check whether a sale is real, there's a free Discount Checker. We're happy to answer questions about the pipeline in the comments.

Top comments (1)