DEV Community

Cover image for Shopify App Store Ads: what 310,000 impressions say about Relevance, bids and dead keywords
Arvio
Arvio

Posted on Edited on Originally published at arvio.a.xyz

Shopify App Store Ads: what 310,000 impressions say about Relevance, bids and dead keywords

This is a repost. Originally published on the Arvio blog: https://arvio.a.xyz/blog/shopify-app-store-ads-benchmarks. The canonical URL points back there.

Key takeaways

  • The keywords Shopify calls "Very high" relevance are the ones most likely to never serve. 65.7% of ours got zero impressions in a year, against 20–27% in every other band. They were 58.4% of our list and took 6.4% of the impressions. Same caveat as every silence figure here: we bid under the suggested range 92.5% of the time, so some of it is ours.
  • Shopify's Relevance label is a click-through predictor, not a quality score. Click-through spans 3.6× across the five bands; click-to-install moves far less — 30.0% across the top two bands against 24.9% across the bottom three.
  • A "Very low" keyword costs 1.83× per install, and 1.59× of that is simply a higher price per click, not worse traffic. The merchants who arrive from a badly-matched keyword install at nearly the same rate; there are just fewer of them and each one costs more to get.
  • The sharpest divider between a live keyword and a dead one, in both our accounts, is whether Shopify shows a bid range at all — not how low the range is. Keywords with a range got 8.95× the impressions per slot over a year, and 3.18× over a week in the other account. Two groups that differ; we cannot show the range causes it.
  • Nearly half of a keyword list can be inert: 49.0% had zero impressions over a full year — though on the keywords with a suggested range we bid under the bottom of it 92.5% of the time, so read that as "produced nothing at our bids".

These numbers came out of our own ad account

We ran the campaigns behind every figure here. The app they were buying installs for is Arvio, an AI store operator that works on the merchant's catalogue rather than their ad spend.

See Arvio on the Shopify App Store →

1. Relevance predicts clicks. It barely predicts installs.

Dataset A, 310,520 impressions, split by the Relevance label Shopify assigns each keyword:

Relevance Impressions Clicks CTR Installs Click → install Cost per install (indexed)
Very high 19,982 224 1.121% 64 28.6% 1.00×
High 41,815 339 0.811% 105 31.0% 1.20×
Medium 74,207 487 0.656% 124 25.5% 1.47×
Low 99,321 561 0.565% 137 24.4% 1.60×
Very low 75,195 233 0.310% 58 24.9% 1.83×
All 310,520 1,844 0.594% 488 26.5%

The CTR column is perfectly monotonic across all five bands, with a 3.6× spread from top to bottom. As a predictor of whether a merchant clicks your card, the Relevance label works.

Now read the next column. Click-to-install is 31.0% for "High" and 24.9% for "Very low", and it is not monotonic — "High" beats "Very high", on 339 and 224 clicks respectively, which is well inside the noise. Collapsed to halves it does carry a signal, just a small one: 30.0% across the top two bands (169 installs on 563 clicks) against 24.9% across the bottom three (319 on 1,281). So Relevance is not silent about install rate; it is 3.6× louder about click-through.

Where the cost difference actually comes from. A "Very low" keyword costs 1.83× per install compared with "Very high". In a pay-per-click auction, cost per install is price per click divided by click-to-install, so that 1.83× has to decompose into exactly those two — and the install-rate half accounts for only 1.15× of it. The other 1.59× is what we paid per click.

Only three columns are new here; the rest are the table above.

Relevance Cost per click (indexed) Bid we had set (median) Bid on the keywords that got a click (median)
Very high 1.00× $1.10 $1.10
High 1.30× $1.10 $1.20
Medium 1.31× $1.10 $1.30
Low 1.37× $1.10 $1.30
Very low 1.59× $1.10 $1.40

The third column stops the obvious explanation and the fourth supplies the real one. We did not bid more on irrelevant keywords: our set bid is a flat $1.10 median in every band. But among the keywords that actually got a click, the bid we happened to have set climbs steadily as relevance falls. At a flat bid, low-relevance keywords mostly never clear the auction; the ones that do are the handful we were over-bidding. So the 1.59× is not Shopify charging more for irrelevant traffic — it is that cheap traffic in those bands did not exist for us to buy. A selection effect, not a price list.

We can show you that from our own account and no further. Whether the same shape holds at a different bid level is exactly the experiment we have not run.

If you have been treating Relevance as a proxy for traffic quality, it is worth separating those two ideas. In this data it is mostly a click-through predictor — loudly so — and only faintly a quality one.

2. The auction sends volume away from the keywords Shopify calls relevant

This is the part we did not expect. The same dataset, counted by keyword instead of by impression:

Relevance Keywords Share of list Impressions Share of impressions
Very high 4,453 58.4% 19,982 6.4%
High 997 13.1% 41,815 13.5%
Medium 1,168 15.3% 74,207 23.9%
Low 785 10.3% 99,321 32.0%
Very low 217 2.8% 75,195 24.2%

Nearly six in ten of our keywords were labelled "Very high" relevance by Shopify, and together they captured 6.4% of the impressions the account received. At the other end, 217 keywords — under 3% of the list — pulled almost a quarter of all impressions.

We can't see the auction, so we can't tell you the mechanism with confidence. The shape is consistent with the obvious explanation: a keyword being highly relevant to your app says nothing about how many merchants type it, and the terms that describe your product precisely tend to be the terms nobody searches. Relevance is a match score, not a demand signal, and it is easy to read it as both.

The operational version: a keyword list that looks excellent in the Relevance column can still be starved of volume, and the dashboard will not flag this. Impression share by relevance band is not a view Shopify gives you. You have to build it.

3. Half of the keyword slots never serve

Dataset A, over a full year: 3,736 of 7,620 keywords — 49.0% — received zero impressions.

Dataset B, over 7 days: 5,819 of 7,609 slots — 76.5% — received zero impressions.

Those two figures are consistent with each other rather than in conflict; a keyword that serves rarely will show zero in a short window and non-zero in a long one. Neither window makes the dead half of the list look alive.

But the aggregate hides the interesting part. Split those 3,736 silent keywords by the label Shopify gave them and the dead half is not spread evenly — it is concentrated in exactly the band you would least expect:

Relevance Keywords Zero impressions in a year
Very high 4,453 65.7%
High 997 24.8%
Medium 1,168 26.2%
Low 785 27.1%
Very low 217 20.3%

Two thirds of the keywords Shopify rated most relevant to our app never served once. They are 58.4% of the list, so they also supply 78% of all the dead slots. This is the mechanism behind section 2: "Very high" does not take a small share of impressions because it is being outbid on live auctions — it takes a small share because most of those auctions never happen.

One concession, made here rather than buried at the end. On the keywords where Shopify offered a bid range, our own bid sat below the bottom of it 92.5% of the time (2,605 of 2,815, same dataset, same year). So an unknown share of these silent slots is not the channel refusing to serve them — it is us not paying to enter the auction. We can't split the two apart with observational data. Read "dead slot" as "produced nothing for us at our bids", not as "has no demand" — and read the takeaway at the top of this page the same way.

4. The only bid signal that separates live from dead is whether a range exists

Shopify shows a suggested bid next to each keyword. Sometimes it is a range ($16.00 – $75.00+), sometimes a single flat figure. We assumed for months that the useful signal was how expensive the suggestion was, and that cheap suggestions meant available inventory.

That was wrong. The useful signal is whether there is a range at all.

The bigger of the two accounts first — Dataset A, 7,620 keywords over a full year:

Bid suggestion Keywords Impressions per slot Never served at all
Shows a range 2,815 92.64 25.4%
Shows a flat figure 4,805 10.35 62.9%
Ratio 8.95×

A keyword with a flat suggestion was two and a half times as likely to be completely inert over a year. The seven-day account points the same way on a much smaller base:

Bid suggestion Slots Impressions per slot Clicks per slot
Shows a range 4,800 1.53 0.00813
Shows a flat figure 2,809 0.48 0.00107
Ratio 3.18× (95% CI 3.00–3.37) 7.6× (95% CI 2.4–24.6)

And the flat suggestions are barely a price at all: 93.1% of them (2,615 of 2,809) are exactly $1.00, which is the floor. Our reading is that a flat $1.00 means Shopify has no auction data for that keyword — nobody is bidding on it because nobody is searching it — and the floor is what it shows when it has nothing to show.

Same direction, twice, on unrelated accounts — 8.95× and 3.18×. Those magnitudes are three times apart and we would not average them or call one a replication of the other. What survives both is the sign and the size of the gap being large, on two apps in two categories over two very different windows. That is the strongest thing in this post, and it is still one advertiser.

A note on that clicks column, because the confidence interval is doing real work. The whole 7-day window contains only 42 clicks. The direction is not in doubt: if clicks were distributed in proportion to slot counts, the flat group should have received 15.5 of those 42 clicks; it received 3 (exact binomial, one-sided, p = 1.0 × 10⁻⁵). But the magnitude — "7.6×" — has a 95% confidence interval running from 2.4× to 24.6×, and anyone quoting 7.6× as a fact is over-reading it. Prefer the impressions-per-slot figures. We publish a 95% interval on the 3.18× because it is what the arithmetic gives, but impressions in this account are heavily concentrated — 2.8% of keywords take 24.2% of them — and clustered counts make any such interval narrower than the uncertainty really is. The replication across two accounts is doing more work here than the interval is.

On keywords where Shopify did show a range, our bid was below the bottom of the suggested range 84.5% of the time in Dataset B (4,058 of 4,800) and 92.5% in Dataset A. We were not outbid; we were not in the auction.

The product on the other side of this spend

Arvio reads a live Shopify catalogue, ranks what is worth fixing, drafts each change against the real data and holds it for approval. Nothing in this article was measured with it.

Install Arvio on the Shopify App Store →

What we'd do differently, if we were starting this account again

Not advice, just the four things we would have set up on day one if we had known:

  1. Build the impression-share-by-relevance-band view on day one. It is the table in section 2, and the dashboard does not have it — which is why nobody looks at it.
  2. Use the presence of a bid range as the first filter on a keyword list, before looking at the suggested amount. A flat $1.00 is not a bargain, it's an empty auction.
  3. Stop reading Relevance as traffic quality. It's a click-through predictor. Judge quality on click-to-install, which in this account was roughly flat across every band.
  4. Count dead slots monthly. Half a keyword list can go quiet without any dashboard changing colour.

The first two are a groupby and a division — the FAQ at the bottom has the recipe, and you don't need anything from us to run them. If you'd rather compare notes than build it, mail us at hi@a.xyz with your band-level table and we'll send ours back: we're the team behind Arvio, and we'd take a second account's numbers over another blog post any day.

Method, and why you should be sceptical of this post

We build Arvio, an AI store operator for Shopify — it audits a store's SEO, product content, prices and stock, drafts the fixes, and holds every change for the merchant to approve. Dataset B below is Arvio's own ad account; Dataset A is our sister app, AiLead. We buy this inventory ourselves, so we have an obvious interest in how you think about the channel, and you should weigh the post accordingly.

What we can offer instead of trust: the exact method, the exact sample, the things that would break the conclusions, and the arithmetic. Every number below comes from keyword-level exports out of the Shopify Partners ad dashboard. No traffic or spend is modelled, estimated or extrapolated — the counts are counts. Where we do compute something on top of them (the indexed cost columns, the confidence intervals, one binomial test) it is labelled as such, and the arithmetic is in the open.

We're publishing it because there is no third-party benchmark for this channel. If you run Google Ads or Meta you can look up what a normal CTR is. For Shopify App Store Ads there is nothing — every advertiser is calibrating against their own account and no one else's. This is one account's worth of ground truth. It is not the industry, but it is more than zero.

Sample. Two datasets, from two apps in two different categories, both ours:

Dataset A Dataset B
App AiLead (sales chatbot) Arvio (store operations)
Rows 7,620 keywords 7,609 keyword slots
Window Last 365 days (Jun 2025 – Jun 2026) Last 7 days (Aug 2026)
Volume 310,520 impressions · 1,844 clicks · 488 installs 8,673 impressions · 42 clicks
Match type 7,478 Broad · 142 Exact 7,599 Broad · 10 Exact
Geo Not split — all rows are All Same

The windows are different and we do not merge them. Dataset A answers "what does Relevance do to performance". Dataset B answers "which keyword slots serve at all". Any sentence below is about one dataset or the other, never both.

Selection bias, stated plainly. These are keywords we chose to buy, mined from our own search-term reports. This is therefore the distribution of a mid-size advertiser's keyword list — not the distribution of App Store search demand. If you want to know what merchants search for, this dataset cannot tell you.

How the rates are computed. CTR per band is impression-weighted (total clicks ÷ total impressions in the band), not the mean of per-keyword CTRs. That choice matters, so we checked it: the unweighted per-keyword mean gives 1.088% / 0.851% / 0.646% / 0.491% / 0.301% — same order, same monotonicity. The finding is not an artefact of weighting.

What we excluded, and what happened when we did. Both datasets are overwhelmingly Broad match, with a small Exact tail. Dropping the Exact rows entirely moves the Dataset A CTR curve from 1.121 / 0.811 / 0.656 / 0.565 / 0.310 to 1.121 / 0.808 / 0.656 / 0.565 / 0.310 — one digit in one band. We kept them.

Right-censoring. In Dataset B, 7.0% (536/7,609) of bid suggestions are capped at the display ceiling $75.00+. Any average of suggested bids is therefore biased low, and we don't publish one.

Timing caveat. Bid suggestions are read at export time. They are not locked to the metric window, so a suggestion shown next to a 7-day impression count was not necessarily the suggestion in force during those 7 days.

Account size, so you can calibrate. This is a small advertiser: annual spend on this channel is in the low four figures, not five or six. Every rate below should be read as a small account's numbers — we have no way to know whether they hold at ten or a hundred times the budget, and we'd be surprised if all of them did.

What we deliberately don't publish. Absolute spend and absolute cost per install. Those are commercial. Everything cost-related below is expressed as a ratio within the same dataset, which preserves every conclusion and leaks nothing. Given the impression, click and install counts above, anyone determined to estimate our costs can get close; the ratios are what we're standing behind.

What we still don't know

  • One advertiser, two apps, no control. We cannot separate "this is how the channel behaves" from "this is how our two apps behave in their two categories."
  • Installs are Shopify-attributed. We cannot verify them independently at merchant level, so the install column inherits whatever Shopify's attribution does.
  • No geo split. Both exports report All. Country-level effects, which are large in every other channel we run, are invisible here.
  • We haven't tested the causal version. Everything above is observational. We have not taken a set of keywords and moved them between bid levels to see what happens — so read section 4 as "these two groups differ", not "showing a range causes impressions".
  • Bid suggestions aren't time-aligned with the metric window (see Method).

If you run App Store Ads and your numbers disagree with ours, we'd genuinely like to know — a second account would roughly double the amount of public data on this channel.

FAQ

What is a good CTR for Shopify App Store Ads?
In this account, across 310,520 impressions over a year, the blended CTR was 0.594%. By relevance band it ran from 1.121% down to 0.310%. We'd caution against treating one account as a benchmark, which is exactly why we published the band-level split rather than a single number.

Does a higher Relevance score get me cheaper installs?
Cheaper, yes — 1.00× versus 1.83× per install between the top and bottom bands in this data. But almost none of that is traffic quality: click-to-install runs 30.0% across the top two bands and 24.9% across the bottom three, which accounts for only 1.15× of it. The other 1.59× is straightforwardly what we paid per click.

Why do my "Very high" relevance keywords get no impressions?
Ours mostly didn't either, and it is the sharpest split in the whole dataset: 65.7% of our "Very high" keywords got zero impressions in a year, against 20–27% in every other band. Relevance measures how well a keyword matches your app, not how many merchants search it — a perfect match for a phrase nobody types is still a perfect match.

What does a flat $1.00 bid suggestion mean?
In our data it is overwhelmingly the floor shown when there is no auction to price — 93.1% of all flat suggestions were exactly $1.00, and those slots served at roughly a third the rate of keywords with a range.

Should I bid above the suggested range?
We can't answer that from observational data. We can say we were below the bottom of the range on keywords that had one 84.5% of the time in one account and 92.5% in the other — which is worth knowing before you conclude that a keyword "doesn't work".

How many keywords should I add?
This data doesn't support an answer, but it does undercut the premise: 49% of ours produced nothing in a year — and we bid under the suggested range 92.5% of the time on the keywords that had one, so some of that silence is ours. Either way, adding slots is not the same as adding reach.

Is Broad or Exact better?
We can't tell you — our list is 98% Broad in both datasets, so we have no meaningful Exact comparison. Anyone claiming a Broad-vs-Exact benchmark for this channel should be asked for their Exact sample size.

Can I reproduce this?
Yes, from your own account. Export keyword-level performance from the Shopify Partners ad dashboard with the Relevance, Impressions, Clicks, Installs and Spend columns, then per band compute four things:

  • sum(Clicks) / sum(Impressions) — the CTR curve
  • sum(Installs) / sum(Clicks) — click-to-install
  • sum(Spend) / sum(Clicks), indexed to your best band — the cost-per-click column that turned out to carry the whole cost story
  • count(Impressions == 0) / count(*) — the dead rate, which is the one nobody looks at

Aggregate the sums per band before dividing; averaging per-keyword rates weights a keyword with 3 impressions the same as one with 30,000. If it doesn't replicate on your account, that's a more interesting result than this post.


Written by the team behind Arvio: AI Store Operator — the AI agent that fixes Shopify store SEO, product content, prices and stock in bulk, with every change waiting for your approval. We're a small app and we buy this ad inventory ourselves; treat this post as one advertiser opening its books, not as an industry study.

Arvio: AI Store Operator — install it on the Shopify App Store. The app whose ad account produced every number above.


Originally published at https://arvio.a.xyz/blog/shopify-app-store-ads-benchmarks. More Shopify bulk-editing writeups are on the Arvio blog.

Top comments (0)