DEV Community

Cover image for AI agent for ecommerce: there are two different jobs, and the search results only show you one
Arvio
Arvio

Posted on Originally published at arvio.a.xyz

AI agent for ecommerce: there are two different jobs, and the search results only show you one

This is a repost. Originally published on the Arvio blog: https://arvio.a.xyz/blog/ai-agent-for-ecommerce. The canonical URL points back there.

Published September 27, 2026 by Adot Technologies Inc, the team behind Arvio. Vendor figures were read on September 27, 2026 from the vendors' own homepages; catalogue figures come from public storefront endpoints any browser can fetch. Method, sample and what we could not read are all below.

What this post is and isn't. It is not a ranked list of the best AI agents, and it is not a feature scorecard — we did not sign up for any of these products or run them against each other. What we did is narrower and checkable: we opened the homepage of every site on Google's first page for this search and counted which kind of work each one advertises. That tells you which aisle you are standing in. It does not tell you what any of these products can actually do. There is a section near the end on when the right answer is a support tool and not us.

The two-minute version: which aisle are you in?

Before you compare anything, answer one question about your own store. Write your answer down, then read the table.

In the last month, what cost you the most — questions you didn't answer, or products you didn't fix?

If the thing that hurts is… The agent you want does… What it reads Homepages on the first page advertising that work
Shoppers asking "where's my order", "does this fit", "can I return it" and nobody replying for 9 hours Conversations: inbox, live chat, tickets, handoff to a human Your help docs, order records, the chat transcript 7 of the 9 we could read
A size that sold out three weeks ago and is still on the page as available Records: the product, variant, price and inventory objects themselves Your catalogue, one product row at a time 0 of the 9
Both Two tools. These do not substitute for each other — —

That zero is the finding. Not one of the nine readable homepages used any of eighteen phrases about editing a product, updating inventory, adjusting a price, managing a catalogue or running a campaign — including the one-word needle seo, which is the easiest phrase in that list to hit by accident.

The other two of the nine aren't in the conversation column either: one is a platform for building your own agents, and one matched neither list. Details and the full table are two sections down.

If your problem is the first row, a catalogue tool will not touch it — install a support agent and stop reading comparison posts. If it's the second row, the conversation tools in the first row are not a smaller version of what you need; they are pointed at a different object.

Check the second row on your own store, in about a minute

This is the part nobody is selling you, so here is how to look at it yourself. It needs no install and no command-line tools.

Open this in a browser tab, replacing the domain with your own storefront:

https://your-store.com/products.json?limit=250
Enter fullscreen mode Exit fullscreen mode

Most Shopify storefronts serve that publicly. You will get a wall of JSON. Don't read it — search it.

Then press Ctrl-F (⌘F on a Mac) and search for this exact string — the quotes matter, and there is no space after the colon:

"available":false
Enter fullscreen mode Exit fullscreen mode

The browser's find bar will tell you how many times it occurs. That count is the number of variants currently out of stock across the products on that page — every size, colour and bundle option that has gone to zero. For comparison, search "available":true the same way.

Two things about that URL before you read anything into the number. limit=250 is a hard ceiling: if you have more than 250 products you are only looking at the first 250, and you need &page=2, &page=3 and so on to see the rest. And the raw count scales with catalogue size, so it means little on its own — on two large public catalogues we searched while checking these instructions, 250 products returned 602 and 1,973 out-of-stock variants. Those are big, healthy stores. The number is a prompt, not a verdict.

What the count means, and what it doesn't. A high false count on its own is not a problem — a product where every variant is false is a fully sold-out product, which is a different (and more visible) issue we have written about separately. The state worth hunting is the mixed one: a product where some variants are false and some are true, because that product is still live, still looks available, and quietly cannot fulfil the size the shopper wants. Ctrl-F alone cannot separate the two, so treat the number as a prompt to go look, not as a measurement.

To see an actual example, don't try to scroll — the file is usually one enormous single line, so there is nothing to scroll up through. Instead search for "handle":, which appears exactly once per product. Step through those hits and open any handle on your storefront as your-store.com/products/THE-HANDLE, then see whether the page tells a shopper anything is missing. (Searching "title": does not work for this: variant names use the same key, so it fires several times per product.)

If you want the exact mixed-state percentage rather than a prompt, that needs a script rather than a find bar — it has to group variants by their parent product and count the products where at least one variant is false and at least one is true. That is the measure behind the 15.3% below, and it is exactly the kind of counting nobody does by hand.

Two jobs, two tools

Arvio is an AI store operator — it works on the record side of the store, not the conversation side. What it does and does not touch is spelled out on our App Store listing rather than here, because this post is about the split, not about us.

See Arvio on the Shopify App Store →

What we actually measured, and what it does not prove

Read this before the table, because the table is easy to over-read.

On 27 September 2026 at 11:56 we took the ten domains ranking on Google's first page for ai agent for ecommerce, opened each homepage in a real browser, and pulled document.body.innerText. Then we searched that text for three fixed lists of phrases: 18 describing work on store records (edit product, update inventory, adjust price, manage catalog, bulk edit, seo, ad campaign and eleven more), 16 describing work on customer conversations (live chat, support ticket, help desk, resolution rate, shared inbox, abandoned cart, product recommendation and nine more), and 10 describing tooling for building your own agent (no-code, agent builder, drag-and-drop, build your own and six more).

This measures what each company advertises on its own front page. It does not measure what the product can do. A homepage is a marketing page. A vendor whose homepage never mentions inventory may still have an inventory feature two clicks away, and a vendor matching customer service twelve times may do a great deal more than chat. Everything below is a claim about copy, and it is phrased that way on purpose. If you want to know whether one of these products edits your catalogue, ask them — this post is not evidence either way.

Nine of ten, not ten of ten. One of the ten domains did not return a page to us on this run: the browser reported a closed connection and a plain HTTPS retry failed at the TLS layer. We could not read it, so it is excluded from every count and percentage below — it is not counted as a zero in any column. Nine is the denominator throughout. We're leaving it unnamed because a domain failing to load on one afternoon says something about that afternoon, not about the company. A blank is not a zero, and "not readable" is its own bucket rather than a rounding convenience.

How a page gets sorted into a bucket. Zero record-work phrases on all nine, so that column never decided anything. Between the other two, the rule is relative weight, not mere presence: a page is called a builder only if it matched more builder phrases than conversation phrases. That distinction matters and we got it wrong the first time — an early version flagged any page containing a builder phrase, which classified Tidio as a builder platform on the strength of two matches for its own drag-and-drop automation feature, against nine matches about support. A feature named "visual automation builder" is not a company describing itself as a builder platform. Under the relative rule it sorts as a conversation tool, which is plainly what its own page title says it is.

What the nine readable homepages advertise

Homepage, as read on 27 September 2026 Its own page title Record-work phrases Conversation phrases Builder phrases Bucket
delight.ai Delight AI: AI concierge for stellar customer experiences 0 2 0 conversations
tidio.com AI Customer Service That Turns Chats Into Revenue | Tidio 0 9 2 conversations
mindstudio.ai Build powerful AI agents | MindStudio 0 0 6 builder
layer3labs.io Layer3Labs - OpenAI, Claude, Gemini for Business 0 2 0 conversations
yourgpt.ai YourGPT | AI Agent Platform for Support, Sales & Operations 0 6 3 conversations
molin.ai Molin AI - Cut your customer support by 80% with AI 0 8 0 conversations
techmonk.io TechMonk | Autonomous AI Agents for Customer Engagement 0 3 3 conversations
fin.ai Fin. The highest performing Customer Agent 0 5 0 conversations
triplewhale.com Triple Whale — Make waves. 0 0 0 matched nothing
(one domain) — not readable not readable not readable excluded

0 of 9 record work. 7 of 9 conversations. 1 builder. 1 that matched none of the 44 phrases.

One row deserves a note rather than a silent pass. techmonk.io tied, 3 and 3 — the rule needs more builder phrases than conversation phrases to call something a builder, so a tie stays in the conversation column. Read that row as the weakest classification in the table.

Several of these pages do describe touching store objects; they just describe it from inside the conversation. Molin's homepage lists "Order updates — Update your customers on their orders" and "Product search — Recommend better products to customers." Tidio's lists "Save abandoned carts" and "Preview carts, check order history." YourGPT's says "A chatbot answers. An agent finishes the job." Note the shape: each one is something done for a shopper who is currently talking to you. None of them is something done to a product page at 2am when nobody is talking to you at all.

The one that matched nothing at all

triplewhale.com returned a full page — 7,844 characters of text, so this is not a short page we failed to read — and none of our 44 phrases appeared anywhere in it. Its title is "Triple Whale — Make waves."

That bucket means our measure did not fit, not that the product does nothing. A page can sell something real in vocabulary our three lists were never built to catch — analytics and attribution, for instance, are neither conversations nor catalogue edits, and no phrase in any of our lists would find them. It stays where the data put it. The same goes for mindstudio.ai in the row above, except there the evidence is positive rather than absent: zero conversation phrases against six builder ones, under a page title that says "Build powerful AI agents". It is a place to build an agent, not an agent for a store.

What we found in 41 real catalogues

The split above is about how a market describes itself. This section is about whether the record-side job is real work, and it comes from different data.

Sample. 41 live Shopify storefronts, 20,604 products, read on 26 August 2026 from /products.json, taken at random from a public list of stores whose owners had posted their own URL on the Shopify Community's Store Feedback board. Everything here came from endpoints any browser can fetch. Zero stores hit the pagination limit on that run, so nothing large was silently truncated.

Whether it looks like you. Stores that ask for feedback skew newer and smaller, and the biggest distortion is survivorship: domains that no longer serve a catalogue were dropped at the frame stage, so what is left is the set that is still trading.

The silent one: a product that is live, and partly gone

Measure, across 41 storefronts Value
Products with more than one variant 7,088 (34.4% of all 20,604)
Of those, products with some variants sold out but not all 1,088 — 15.3%
Same figure across all products, multi-variant or not 5.3%
Stores with at least one such product 28 of the 39 that had any multi-variant products
Median store's rate, within its multi-variant products 6.2%

Pooled 15.3% against a median store of 6.2% — a ratio of 2.48 — means a minority of catalogues carry most of it. Do not read 15.3% as "this is what your store looks like." Read it as: this is common enough that it is worth spending one minute checking, and if you are in the heavy group you are much worse than 15.3%.

We also checked whether the high-rate stores were just tiny stores with three multi-variant products each, where one sold-out size makes a big percentage. They were not: of the stores at or above 15%, 11 had ten or more multi-variant products and only 3 had fewer than ten.

Four things this data cannot tell you, and we are not going to guess at them:

  • How many units are left. The public endpoint gives each variant a true/false available flag and no quantity. No "fewer than N in stock" claim can be computed from this, ours included.
  • What the product page shows the shopper. Greyed-out size, or a full sold-out banner? We read the records, not the rendered pages.
  • Whether those variants were discontinued on purpose. No store told us why anything went to zero, and a deliberately retired colourway looks identical here to an oversight.
  • How fast anything is restocked. This is a single-day snapshot. There is no before-and-after pair, so nothing about speed can be said.

That first limitation matters for the argument: we can show you that a variant is gone, not that anyone is losing money over it.

The other record-side finding: imported products are thinner

Same 41 stores. We grouped products by store and creation date, and treated 20 or more products created by one store on one day as a bulk creation batch — 132 such batches across 20 of the 41 stores, accounting for 53.5% of all products.

Field, as returned by /products.json Products in a batch of 20+ (n=11,014) Everything else (n=9,590)
Empty description 4.4% 0.7%
Thin description 29.7% 11.9%
No product_type 38.0% 35.0%
No tags 21.6% 17.9%
One image or none 46.4% 21.0%

The caveat that kills the obvious headline. It is tempting to write "the bigger the import, the dirtier the data," and the data does not support it. We ran the same split at three thresholds, and also cut the products into non-overlapping size bands. In those bands the 20–49 band is the cleanest in the whole sample — 1.9% thin descriptions, 7.1% without tags, better than products created one at a time. Almost the entire effect sits in the 50-and-over band (38.4% thin, 56.4% with one image or none). So the honest statement is: very large batches are much thinner; medium ones are not, and are sometimes better than hand-built products. Our batch rule is also date-based, so a merchant who spent a careful Tuesday adding 25 products by hand gets counted as a batch.

Also, two fields that look damning are not: product_type was blank on 36.6% of all 20,604 products and tags on 19.9% — but those gaps are nearly as large outside the batches as inside. Those are not import damage. That is just a field nobody fills in.

How this sits next to Klaviyo, if you already run it

We are asked this often enough to answer it plainly: Klaviyo is not on the other side of a comparison table from us. The two tools read different halves of the same store, which is why "which one should I buy" is usually the wrong question about them.

What it reads What it acts on
Klaviyo's AI Customer and event data: who bought, who browsed, who lapsed Messages — the email or SMS that goes out, who gets it, when
Arvio Product records: description, type, tags, images, price, variant availability The records themselves, each change held for your approval

The seam between them is the thing worth setting up. A back-in-stock flow is only as good as the availability flag behind it: if a variant flipped and the record never moved, the flow has nothing to fire on. A winback campaign pointed at a product whose description is two sentences is sending traffic to a page that cannot close. Fix the records, then let the messaging tool do what it is good at. Neither one replaces the other.

When you should stay where you are

Most posts on this search term are written by a vendor, and the useful thing a vendor can do is tell you when not to buy from them.

If what hurts is that buyers can't reach anyone, install a support agent, not us. Concretely: your inbox has unanswered messages older than a day; the same five questions arrive every week about shipping times, sizing and returns; you answer chats yourself between 9pm and 11pm; someone asked pre-purchase whether an item fits, got no reply, and didn't buy. That is a staffing problem with a software answer, and the answer is a tool built for conversations — Tidio and Fin are two of the names on the first page for this search, and their own homepages say plainly which job they do. Arvio will not answer a single one of those messages. It cannot see the message.

If the agent you are weighing is Shopify's own, we ran the same split on it in Shopify AI agent vs Sidekick.

A test that separates them cleanly. Count what you lost last month in two buckets: revenue lost because nobody replied, versus revenue lost because a page was wrong, thin or quietly unavailable. Buy for the bigger bucket. For a lot of stores under a thousand orders a month, the bigger bucket is the first one, and it is not close.

Two more cases where we are the wrong answer. If your catalogue is under about thirty products, you can do the record work yourself in an evening and an agent is overhead — nothing above applies to you at that size. And if what you need is a custom workflow wired across your own internal tools, you want a builder platform — which is a real category, and one of the nine homepages we read is selling exactly that. A packaged store agent, ours included, will be too opinionated about what it touches.

What we did not test

We did not install, trial or use any of the nine products. We did not compare quality, resolution rates or pricing, and nothing here says any of these products is better or worse than any other — only which job each one puts on its own front page. We also did not test whether any of them edit catalogue records; the absence of those phrases from a homepage is not evidence of absence in the product. The catalogue figures come from a convenience frame read on one day through one public endpoint, and they describe those 41 stores, not Shopify.

Knowing which aisle you're in doesn't fix the shelf

If the answer was the record side, somebody still has to open each product and type. Arvio sits on that side of the split; the listing is the place to judge whether it fits your store.

Install Arvio on the Shopify App Store →

FAQ

What is an AI agent for ecommerce?

The phrase currently covers at least two unrelated jobs, and the search results only show one of them. On the nine first-page homepages we could read on 27 September 2026, seven advertise work on customer conversations — chat, tickets, inbox, handoff. Zero advertise work on store records. When someone says "AI agent for ecommerce" without qualifying it, they almost always mean the first one, because that is what the market looks like.

Are all ecommerce AI agents just chatbots?

No, and our data can't support that claim — it is about homepage copy, not products. Of nine readable homepages, seven advertise conversation work, one is a platform for building your own agents (zero conversation phrases, six builder ones), and one matched none of our 44 phrases at all. A tenth domain would not load for us and is excluded from those counts entirely.

Why is your denominator nine and not ten?

Because one of the ten domains did not return a page on the run we used. The browser reported a closed connection, and the plain HTTPS fallback failed at the TLS layer. Counting it as "zero record-work phrases" would have strengthened our headline and would have been false — we did not read that page, so we know nothing about it. We don't name it, because a single load failure on one afternoon says more about that afternoon than about the company. Not readable is its own bucket.

Does a support agent handle inventory problems?

The ones whose homepages we read describe order lookups, cart recovery and product recommendations inside a conversation. That is different from noticing, with nobody in the chat, that one variant of a live product went to zero three weeks ago. In our 41-store sample that state existed on 15.3% of multi-variant products. We went into why Shopify's own alerting misses it in Shopify does not have a low stock alert. Whether any specific product handles it is a question for that vendor.

Can I check the 15.3% number on my own store?

Not exactly, with a browser alone. The Ctrl-F method near the top counts out-of-stock variants, which is a prompt to investigate rather than the same measure. The 15.3% counts products where at least one variant is unavailable and at least one is still available, which requires grouping variants under their parent — a script, not a find bar. The find-bar count will tell you within a minute whether there is anything to investigate at all.

Does Arvio answer customers?

No — it sits on the record side of the split this whole post is about, and it cannot see a customer message. If you need customer replies, the section above on staying where you are names the category you actually want.

Is Klaviyo a competitor?

No — the two act on different halves of the store. Klaviyo's AI acts on messages; Arvio acts on the product records those messages point at. The failure mode we see is running good campaigns into pages that are thin or quietly out of stock.

One tab, one Ctrl-F

If you read only one thing here: the two aisles are real, they do not overlap, and the first page of results only shows you one of them — nine homepages, eighteen chances each to mention catalogue work, zero mentions.

Open /products.json?limit=250 on your own storefront and search it for "available":false. If you get no hits, your record side is quiet today and a conversation tool is probably the better spend. If you get a screenful, you have just found work that nobody was ever going to file a ticket about.

Arvio: AI Store Operator — see it on the Shopify App Store. We are on the record side of the split above — the listing is where to check whether that is the side your store is losing money on.


Originally published at https://arvio.a.xyz/blog/ai-agent-for-ecommerce. More Shopify bulk-editing writeups are on the Arvio blog.

Top comments (0)