Every tool you are about to buy has two review scores, and they do not agree.
Ahrefs is 4.5 on G2 and 1.8 on Trustpilot. Semrush is 4.5 and 2.0. Klaviyo is 4.6 and under 2. Mailchimp is 4.4 and around 2.
I kept running into this while pricing 20+ marketing tools from their published rate cards, so I went and checked all four properly. The gap is not noise, and it is not the two crowds disagreeing about whether the product is good. They are answering different questions, and once you know which question each site is answering, the gap becomes the single most useful thing on either page.
The numbers
| Tool | G2 | Capterra | Trustpilot | G2 − Trustpilot |
|---|---|---|---|---|
| Ahrefs | 4.5 (675) | 4.7 (583) | 1.8 (306) | 2.7 |
| Klaviyo | 4.6 (1,361) | 4.6 (537) | under 2 (367) | ~3 |
| Semrush | 4.5 (3,434) | 4.6 (2,318) | 2.0 | 2.5 |
| Mailchimp | 4.4 (12,885) | 4.5 (17,626) | ~2 | ~2.4 |
Review counts in brackets. Scores read from each platform in mid-2026.
Four products, four different companies, two different categories — SEO tooling and email — and the same 2.4-to-3-point spread every time. A pattern that consistent is structural, not a run of bad luck.
What each site is actually asking
The two populations are selected differently, and they are asked different questions at different moments.
G2 and Capterra ask "does this tool do the job?" The reviewer is usually mid-evaluation or recently onboarded. They are rating capability, interface, and fit against the alternatives they just compared. Many arrive through a vendor's own review campaign, gift card attached. Nobody in that flow has yet had the experience of a renewal invoice arriving 60% higher than the number they budgeted.
Trustpilot asks "would you recommend this company?" Almost nobody visits Trustpilot mid-evaluation. People go there after something happened to them, and in SaaS that something is nearly always billing: a trial that auto-converted, a plan that jumped a tier, a cancellation flow that took three emails. The population is self-selected and skews negative — that is true across every category on the site, not just software.
So the ratings gap is not the product score minus the product score. It is the capability score minus the billing score. And that framing makes it testable.
The gap tracks the pricing model
If the gap really is a billing signal, then the tools with the widest gaps should be the ones whose real cost is least visible during evaluation. That is exactly what I found — each of these four has a pricing mechanic that does not appear in the comparison table you used to decide.
Ahrefs — the credit meter. Reports, exports, and deep crawls all draw down one credit pool. Seat pricing is predictable because you know how many people you have; credit pricing moves with how hard you work, so a busy month costs more precisely when you are busiest. There is also no free trial in 2026, so you pay a full month before you learn any of this. Advertised entry $29, realistic working plan $129 — a 4.4× multiple. Full breakdown.
Semrush — seats are not included. Every plan ships with exactly one. A three-person team on Guru pays $249.95 plus two seats at $80, so $409.95 a month: 64% above the advertised plan price for the same plan. The Pro-to-Guru step is itself a 78.6% increase, the steepest tier transition in the category, and Pro is deliberately incomplete for professional work.
Klaviyo — billed on stored profiles. Not sends, not engaged subscribers. Every abandoned signup and every one-time customer from 2023 is on the meter monthly. Going from 10,000 to 11,500 profiles — 15% more people — takes the bill from $150 to $225. That is a 50% increase for 15% more list, and it lands exactly where a successful small store ends up. The cliff, charted.
Mailchimp — the contact count includes people who will never open an email. Unsubscribed and bounced records stay billable until you manually archive them, and the same person in two audiences is billed twice. Advertised $13, realistic $135 — 5.2×. Details.
Four wide gaps, four pricing mechanics that are invisible at evaluation time and unavoidable on an invoice. The correlation is the finding.
Where this sits in the wider data
This is not confined to four tools. Across 13 tools with a paid entry tier, the plan you realistically end up on costs 3.9× the advertised entry price — median 3.3×, with Close at 11× ($9 to $99). Monthly billing carries an average 44% premium over annual, and the premium is steepest at the cheap end, where Close charges 111% and Pipedrive 71%. Only 4 of 17 tools have a genuinely usable free tier. Full dataset and method.
A category that routinely prices at 3.9× its own headline number is a category that will generate a 2.5-point ratings gap. The two facts are the same fact.
The honest caveats
I would rather you trust the framing than the decimal places, so:
- Trustpilot is self-selected and negative-skewed everywhere. A 1.8 does not mean 1.8 out of every 5 customers is happy. Use it as a signal about what kind of complaint clusters, not as a satisfaction rate.
- Sample sizes are wildly unequal. Mailchimp has 12,885 G2 reviews against a few hundred on Trustpilot for some of these tools. Do not read small-n scores as precise.
- G2 and Capterra run vendor-incentivised review programmes. That is not fraud, but it does mean the population is partly recruited by the vendor at its happiest moment.
-
Two of the four Trustpilot figures are approximate — hence
under 2and~2rather than fake decimals.
None of that dissolves a 2.5-point spread appearing four times in a row.
How to actually use this
The practical version is short:
- Read G2 for capability and Trustpilot for billing. Stop treating them as competing estimates of the same quantity. They are two different instruments.
- Read the Trustpilot 1-star reviews specifically, and only for mechanism. Ignore the tone; extract the mechanic. "Auto-renewed after the trial", "charged for contacts I deleted", "ran out of credits" — those are pricing-model facts you can verify on the rate card before you buy.
- Price the tier above the one you think you need, plus a seat for everyone who will log in. That is the number the Trustpilot reviewers were surprised by.
- Treat a large gap as a to-do, not a veto. Ahrefs' index is genuinely the best available and Klaviyo's segmentation is genuinely ahead of the field. A wide gap does not mean "bad product". It means "read the billing terms before the trial converts". Set a calendar reminder for day six of any seven-day trial and most of what the low score is made of never happens to you.
The gap is the most honest number on either review site. It just is not the number either site is trying to show you.
I run RealCostLabs, where I rebuild SaaS pricing from published rate cards and price it at usage levels real businesses actually hit. All figures above are from published rate cards and public review platforms in mid-2026; vendors change prices, so check before you buy.

Top comments (0)