DEV Community

gobuy.ai
gobuy.ai

Posted on

How We Built a Fake-Review Detector Using Wilson Confidence Intervals and Bayesian Inference

Can math tell you whether those 4.9-star reviews are real?

Not with certainty — but with confidence, and the difference between the two is where the interesting engineering lives. This post walks through the statistical core of GoBuy's scoring: two layers of classical inference applied to marketplace reviews, what each one is actually for, and the trap hiding in the most obvious approach.

The problem

Amazon lists over 300 million products, and conservative estimates put fake, AI-generated, or incentivized reviews at 30–40% of the total — a share that grows as LLMs make generation free. Existing tools (Fakespot, ReviewMeta) answered this with proprietary black boxes. We wanted the opposite: a system where every number is derivable from first principles, so anyone can audit why a product scores the way it does.

Layer 1: the Wilson score lower bound — sample-size honesty for ranking

The naive average ignores how many reviews it rests on. A product with 10 reviews at 100% positive and a product with 1,000 reviews at 95% positive both "average" near the top — but they are not equally believable.

The Wilson score interval fixes this. For observed positive proportion p over n reviews at 95% confidence (z = 1.96):

lower = (p + z²/2n − z·√(p(1−p)/n + z²/4n²)) / (1 + z²/n)
Enter fullscreen mode Exit fullscreen mode

Worked examples:

  • 480/500 positive (96%) → lower bound 0.939
  • 950/1,000 positive (95%) → lower bound 0.935
  • 10/10 positive (100%) → lower bound 0.72 — a perfect-but-tiny record still leaves real uncertainty

The 10-review product drops below products with hundreds of slightly-less-perfect reviews. That is the point: the Wilson layer is a ranking tool that makes claims honest about their sample size.

One thing this layer is not: a fraud detector. A lower bound on the positive rate measures how liked a product appears to be, not whether the reviews are real — a coordinated burst raises the rate and shrinks the interval, which is exactly what an attacker produces. Fraud detection lives in different signals entirely (velocity, trajectory shape, verified-purchase mismatches — a future post).

Layer 2: the Beta-Binomial model — beliefs that update over time

Wilson gives a snapshot. Reviews accumulate, and we want beliefs that update incrementally.

Treat the true positive rate as a Beta-distributed random variable. Start with Beta(1,1) (uniform — no prior opinion), then each positive review nudges the parameters up:

Beta(α₀ + positives, β₀ + negatives)
Enter fullscreen mode Exit fullscreen mode

480/500 → Beta(481, 21). The 5th percentile of that posterior is 0.9425 — note how close it sits to Wilson's one-sided 0.943 bound. That is not a coincidence: with a flat prior, both layers are functions of the same two counts, and they agree to the third decimal.

So why keep the Bayesian layer at all? Because its value is not in the flat-prior case — it is in what the flat prior can't do:

  • Real priors. A category-level base rate (e.g., "electronics in this price band average 4.1 stars") gives the posterior something to learn from, pulling thin histories toward category reality instead of letting 10 reviews dominate.
  • Sequential updates. New evidence arrives daily; the posterior updates in O(1) without recomputing history.
  • Regime changes. A prior informed by recent behavior makes sudden distribution shifts visible faster.

The honest engineering summary: with Beta(1,1), the Bayesian layer is nearly the same signal as Wilson counted twice — it earns independent weight only when it carries real prior structure and the temporal dimension. That is how it runs in production.

What the composite actually measures

Neither layer measures review authenticity. Together they measure how much the observed rating deserves its number — the evidence-quality backbone of a trust score. Authenticity (burst detection, trajectory anomalies, verified-purchase mismatches) is a separate family of longitudinal signals layered on top, and it deserves its own post.

The composite, briefly:

  • Wilson lower bound — ranking honesty about sample size (~35%)
  • Bayesian posterior with category priors — temporal + prior structure (~30%)
  • Longitudinal authenticity signals — burst, trajectory, provenance (the rest; future post)

Try it

The full corpus runs at gobuy.ai; the gate that consumes it is GPV-1, our open pre-purchase verification spec — one call, check_product_trust, returns the evidence score an agent should see before checkout. The reference implementation is open source.


Corrections: an earlier version stated 0.69 for the 10/10 Wilson example; the correct value is 0.72 (thanks to @arhancanli for catching it). This version also clarifies the division of labor between the two statistical layers, following his equally fair observation that they nearly coincide under a flat prior.

Top comments (2)

Collapse
 
arhancanli profile image
Arhan Canli •

Checked the numbers: 480/500 gives 0.939 and 950/1,000 gives 0.935, matching yours. 10/10 at z=1.96 comes out at 0.72 for me, not 0.69, so worth rechecking that line.

The bigger point is that the two layers are nearly the same quantity. The Wilson 90% one-sided bound (z=1.645) for 480/500 is 0.943, and the Beta(481, 21) 5th percentile is 0.9425. For 10/10 it is 0.787 vs 0.762, and for 950/1,000 0.9374 vs 0.9372. Both are functions of the same two counts, so giving them 35% and 30% of the score is mostly one signal counted twice, and the Beta layer only adds information if you put a real prior on it (a category-level prior instead of Beta(1,1)).

Also, a lower bound on the positive rate says the product is liked, not that the reviews are real. A fake-review burst raises the rate and shrinks the interval, so layer 1 rewards exactly the thing an attacker produces. Did you test the Smart Score on a set of known-manipulated listings, and does it drop below the clean ones?

Collapse
 
gobuy profile image
gobuy.ai •

Arhan — thanks for checking the numbers, and you're right on all three counts.

The 0.69 is wrong. 10/10 at z=1.96 gives 0.7225 — corrected to 0.72 in the article. That line survived an edit where we switched from a one-sided bound to the two-sided interval and the example number didn't move with it. Good catch.

On double-counting: fair, as the article presents it. With a flat Beta(1,1) prior the posterior quantile is nearly identical to the Wilson bound — as your examples show, they agree to the third decimal. The layer separation only earns independent weight when the Bayesian side carries real prior structure (category-level base rates) and the temporal update; against a flat prior it is, as you say, one signal counted twice. The honest framing: Wilson is the ranking backbone; the Bayesian layer exists to absorb priors — if it isn't doing that, it shouldn't get separate weight in the composite.

On "liked ≠ real": the sharpest point, and we agree. A lower bound on the positive rate measures confidence in popularity, not authenticity — and a burst raises the rate while shrinking the interval, so layer 1 rewards exactly what an attacker produces. In the production score, authenticity comes from entirely different layers: velocity/burst anomalies on the review timeline, rating-trajectory shape, and verified-purchase mismatches — longitudinal, distributional signals. Wilson's job is sample-size honesty in ranking (a 10-review 5.0 should lose to a 1,000-review 4.7), never fraud detection. The article blurred that line; fair hit.

We're publishing the manipulation side — burst detection, trajectory analysis, and validation against known-manipulated listings from our observation ledger — as the next posts in this series. Comments like this are exactly the review the method needs.