<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mike Wei</title>
    <description>The latest articles on DEV Community by Mike Wei (@michael_wei_d93a005ecc379).</description>
    <link>https://dev.to/michael_wei_d93a005ecc379</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4019900%2Fc2e2ed27-540a-4bb8-a833-9887f4a49b0c.png</url>
      <title>DEV Community: Mike Wei</title>
      <link>https://dev.to/michael_wei_d93a005ecc379</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/michael_wei_d93a005ecc379"/>
    <language>en</language>
    <item>
      <title>AI now shows up in 1 in 16 eSIM support reviews. Whether customers like it depends on one thing.</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Sun, 27 Sep 2026 05:48:21 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/ai-now-shows-up-in-1-in-16-esim-support-reviews-whether-customers-like-it-depends-on-one-thing-1jo8</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/ai-now-shows-up-in-1-in-16-esim-support-reviews-whether-customers-like-it-depends-on-one-thing-1jo8</guid>
      <description>&lt;p&gt;Three years ago, almost nobody reviewing an eSIM brand mentioned talking to a bot. Now it's routine. In our refreshed eSIM Customer Service Benchmark, the share of support reviews that mention AI, a chatbot or automated replies went from 0.6% in 2023 to 4.7% in 2024 and 6.4% so far in 2026.&lt;/p&gt;

&lt;p&gt;The more useful finding is what doesn't follow from that. Brands with a lot of visible AI in support don't have happier customers, and brands with very little don't either. The presence of AI tells you almost nothing. What it can do tells you a lot.&lt;/p&gt;

&lt;p&gt;Quick background: the benchmark covers the 20 largest eSIM brands on Trustpilot, 377,827 reviews. For this angle we read 9,883 confirmed support reviews, searched them for 17 AI-related terms and checked every hit by reading the review. Disclosure: I run Aissist, which sells AI support agents, so I have a horse in this race. That's partly why I wanted the data to be public.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where AI is most visible
&lt;/h2&gt;

&lt;p&gt;Across the industry, 5.4% of support reviews mention AI. It is most visible at BNESIM (18%, on a small sample of 82 reviews), Airalo (10%), Ubigi (9.2%), and Saily and eSimFLAG (7.6% each). It barely registers at GigSky (0.9%), eSIM.me (1.3%), Simify (1.5%) and aloSIM (1.8%).&lt;/p&gt;

&lt;p&gt;Treat all of these as floors. AI only gets counted when a customer notices it and says so. A bot that quietly fixes your problem in 40 seconds tends to get reviewed as "support was quick," not "the AI was great." And Airalo and Holafly both hit Trustpilot's 400-result cap on the "AI" and "bot" searches, so their counts are understated.&lt;/p&gt;

&lt;h2&gt;
  
  
  More AI doesn't mean happier customers or unhappier customers
&lt;/h2&gt;

&lt;p&gt;Put AI visibility next to how negative a brand's support reviews are, and there's no clean line:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1wwtzd0c0r1m11t9yew1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1wwtzd0c0r1m11t9yew1.png" alt=" " width="759" height="374"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Saily and Ubigi are the clearest pair. Similar AI visibility, and support reviews that are 16% negative at one and 60% at the other. MobiMatter shows the opposite shape: little visible AI, and 59% of its support reviews negative anyway. Whatever separates good support from bad in this category, "has a chatbot" isn't it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What customers complain about when they complain about AI
&lt;/h2&gt;

&lt;p&gt;In the ranking, Bot walls come fourth. The two biggest support failures across the industry are problems support couldn't fix (29 of every 100) and slow or missing replies (28). Those are the exact problems automation is supposed to solve. An AI that answers instantly fixes the second one. An AI that can't act on the account doesn't fix the first one. It just gets to "sorry, I can't help with that" faster.&lt;/p&gt;

&lt;p&gt;The one thing that matters: can it resolve, and can it hand over?&lt;/p&gt;

&lt;p&gt;My read of these reviews is that good and bad AI experiences come down to two questions.&lt;/p&gt;

&lt;p&gt;Can it actually do something? Reissue the eSIM, check whether data is flowing, see that the plan expired early, apply a credit. An assistant that only answers FAQs is deflection, not resolution, and a traveler stranded without data notices the difference quickly.&lt;/p&gt;

&lt;p&gt;When it can't, does a human take over, with context? The failure mode people write one-star reviews about isn't "I talked to a bot." It's "I talked to a bot and there was no way out." A handoff that carries the conversation over, so the customer doesn't repeat themselves, turns an AI miss into a normal support interaction.&lt;/p&gt;

&lt;p&gt;If you're evaluating AI for eSIM support, those are the two questions to ask vendors. Include us.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;AI is now a visible part of eSIM support and growing fast. On its own, it doesn't move customer sentiment in either direction. The brands getting it right use automation to be reachable instantly and to fix things, with a human one step away. The ones getting it wrong built a faster way to say no.&lt;/p&gt;

&lt;p&gt;Full per-brand data and method: &lt;a href="https://aissist.io/industries/esim-customer-service-benchmark" rel="noopener noreferrer"&gt;eSIM Customer Service Benchmark 2026&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>We read 8,300 low-star eSIM reviews. Support is rarely the cause, and often the multiplier.</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Sun, 27 Sep 2026 05:37:27 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/we-read-8300-low-star-esim-reviews-support-is-rarely-the-cause-and-often-the-multiplier-3gd4</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/we-read-8300-low-star-esim-reviews-support-is-rarely-the-cause-and-often-the-multiplier-3gd4</guid>
      <description>&lt;p&gt;If you run support at an eSIM company, here is the uncomfortable finding from our latest benchmark: customer service is the main complaint in only 11% of 1–2★ reviews, but it shows up in 27% of them. Support rarely breaks the trip. It decides whether a broken trip becomes a one-star review.&lt;/p&gt;

&lt;p&gt;Some context. We refreshed the eSIM Customer Service Benchmark this week with Trustpilot data from the 20 largest eSIM brands, 377,827 reviews in total. For this cut, a model read and labeled 8,300 1–2★ reviews posted in the last 12 months, recording both the main complaint and every other issue each reviewer raised. Disclosure up front: I run Aissist, which sells AI customer support software, so weigh my opinions accordingly. The method and every table are public.&lt;/p&gt;

&lt;h2&gt;
  
  
  The product fails first
&lt;/h2&gt;

&lt;p&gt;Connectivity is the main complaint in 48% of low-star reviews and part of the complaint in 56%. Activation, where the eSIM never works at all, is the main complaint in another 18%. Put those together and two in three angry reviewers are angry because the thing they bought didn't do its one job, usually right after landing.&lt;/p&gt;

&lt;p&gt;A typical one, from an Airalo review:&lt;/p&gt;

&lt;p&gt;"We purchased 2 eSIMs for Tanzania and both didn't work 90% of the time, either in mainland or Zanzibar."&lt;/p&gt;

&lt;p&gt;No support team fixes a tower in Zanzibar. That part is a network and product problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then support makes it worse
&lt;/h2&gt;

&lt;p&gt;Here is where it gets interesting. 57% of support complaints sit on top of a product failure. The traveler had a problem, went looking for help, and the help became a second problem. That is the gap between 11% and 27%: support is the headline in a few reviews and the aggravating factor in many more.&lt;/p&gt;

&lt;p&gt;So what actually goes wrong? We labeled 3,464 low-star support complaints by type. Of every 100 support failures:&lt;/p&gt;

&lt;p&gt;29 are support that engaged but couldn't fix the problem&lt;br&gt;
28 are slow or missing replies&lt;br&gt;
14 are refund refusals&lt;br&gt;
13 are a bot with no path to a human&lt;br&gt;
3 are rudeness&lt;/p&gt;

&lt;p&gt;Rudeness at 3 surprised me. Agents in this industry are, by and large, polite. They are just not fast enough, or not empowered enough. eSIM support doesn't have a manners problem. It has a reach-and-resolution problem.&lt;/p&gt;

&lt;p&gt;The timing makes it harsh. A traveler with no data can't wait until tomorrow for an email reply. One Maya Mobile reviewer put it plainly:&lt;/p&gt;

&lt;p&gt;"The customer service is emails so please do this in advance as the replies are so slow and unhelpful."&lt;/p&gt;

&lt;p&gt;"Please contact support in advance" is not advice anyone wants to give about a product you buy for when you're already abroad.&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure mode differs by brand
&lt;/h2&gt;

&lt;p&gt;The mix matters, because each failure mode calls for a different fix.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn52u2pc277vksjqedotm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn52u2pc277vksjqedotm.png" alt=" " width="735" height="294"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A brand losing on reach and a brand losing on resolution can post the same star rating and need completely different roadmaps.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do with this if I ran eSIM support
&lt;/h2&gt;

&lt;p&gt;Tag support tickets by the product failure underneath them. If most of your support complaints ride on connectivity and activation, your best CSAT project may be a faster troubleshooting flow for those two cases, not a new tone guide.&lt;/p&gt;

&lt;p&gt;Treat time-to-first-response abroad as a product feature. For this category, "we'll get back to you within 24 hours" means "after your trip has already gone wrong."&lt;/p&gt;

&lt;p&gt;Let the front line act. "Couldn't fix it" is often "wasn't allowed to fix it." Whoever answers first, human or AI, needs access to provisioning, usage data and refund rules, or the conversation just becomes a slower route to the same dead end.&lt;/p&gt;

&lt;p&gt;If you use a bot, make the handoff instant. Bot walls are 13% of support failures. Automation that answers fast but can't escalate simply moves the frustration earlier.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;Most one-star eSIM reviews start with a network or activation failure. Support can't prevent those, but it decides how bad they get, and it fails mainly on reach and resolution. Fix those two and you improve the quarter of your worst reviews where support is part of the story.&lt;/p&gt;

&lt;p&gt;Full data, per-brand tables and method: &lt;a href="https://aissist.io/industries/esim-customer-service-benchmark" rel="noopener noreferrer"&gt;eSIM Customer Service Benchmark 2026&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>telecom</category>
    </item>
    <item>
      <title>Claude for Commerce: Hype, or a Bomb Dropped Into the Commerce Ecosystem?</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Thu, 03 Sep 2026 17:43:19 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/claude-for-commerce-hype-or-a-bomb-dropped-into-the-commerce-ecosystem-2did</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/claude-for-commerce-hype-or-a-bomb-dropped-into-the-commerce-ecosystem-2did</guid>
      <description>&lt;p&gt;Anthropic is moving beyond models and into vertical solutions. Commerce is one of its first major bets.&lt;/p&gt;

&lt;p&gt;Anthropic recently launched Claude for Commerce.&lt;/p&gt;

&lt;p&gt;What interests me is not just the product itself, but what it signals.&lt;/p&gt;

&lt;p&gt;Anthropic is moving deeper into business solutions instead of simply providing a foundation model. And eCommerce is a very reasonable place to start.&lt;/p&gt;

&lt;p&gt;Commerce has many workflows that AI can handle well:&lt;/p&gt;

&lt;p&gt;Product recommendations&lt;br&gt;
Order tracking&lt;br&gt;
Returns and refunds&lt;br&gt;
Inventory questions&lt;br&gt;
Customer support&lt;br&gt;
Sales and marketing tasks&lt;/p&gt;

&lt;p&gt;Our own eCommerce benchmark (&lt;a href="https://aissist.io/industries/ai-customer-service-benchmark-2026" rel="noopener noreferrer"&gt;https://aissist.io/industries/ai-customer-service-benchmark-2026&lt;/a&gt;) also shows that this is already one of the areas where agentic AI can create meaningful business impact.&lt;/p&gt;

&lt;p&gt;But there is still a big difference between a good demo and a production-ready solution.&lt;/p&gt;

&lt;p&gt;In commerce, AI needs to be more than intelligent. It needs to be reliable, fast, and predictable.&lt;/p&gt;

&lt;p&gt;If an AI recommends slightly different products each time, that's probably fine.&lt;/p&gt;

&lt;p&gt;If it processes the same refund request differently each time, that's a problem.&lt;/p&gt;

&lt;p&gt;And this is where I think the ecosystem becomes interesting.&lt;/p&gt;

&lt;p&gt;Anthropic can provide the intelligence and agent framework. Shopify and other platforms provide commerce infrastructure. Payment companies provide transactions.&lt;/p&gt;

&lt;p&gt;But someone still needs to connect everything together and make sure the system actually works end to end.&lt;/p&gt;

&lt;p&gt;So I don't think Claude for Commerce immediately replaces existing commerce platforms or AI vendors.&lt;/p&gt;

&lt;p&gt;Instead, it could push the whole market forward—and force everyone to rethink where they add value.&lt;/p&gt;

&lt;p&gt;The bigger question is:&lt;/p&gt;

&lt;p&gt;Will Anthropic stay as the intelligence layer, or eventually move further up the stack and own more of the solution?&lt;/p&gt;

&lt;p&gt;And for companies already building AI for commerce:&lt;/p&gt;

&lt;p&gt;Is Claude for Commerce a threat, or does it make the market much bigger?&lt;/p&gt;

&lt;p&gt;Curious what everyone thinks.&lt;/p&gt;

&lt;p&gt;Full analysis:&lt;br&gt;
&lt;a href="https://aissist.io/insights/claude-for-commerce" rel="noopener noreferrer"&gt;https://aissist.io/insights/claude-for-commerce&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Your AI resolution rate went up and CSAT went down. That's not a tradeoff — it's a measurement bug.</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Wed, 19 Aug 2026 18:01:50 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/your-ai-resolution-rate-went-up-and-csat-went-down-thats-not-a-tradeoff-its-a-measurement-bug-2d2k</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/your-ai-resolution-rate-went-up-and-csat-went-down-thats-not-a-tradeoff-its-a-measurement-bug-2d2k</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: I work at Aissist.io, which builds AI agents for customer support. This is adapted from a &lt;a href="https://aissist.io/industries/resolution-csat-tradeoff" rel="noopener noreferrer"&gt;longer piece on our site&lt;/a&gt;. The argument stands on its own; judge it on the logic.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Every support team that deploys an AI agent eventually hits the same wall. Resolution rate climbs from 40% to 60% to 75%, the dashboard turns green, and then CSAT starts sliding. The conclusion looks obvious: automation and satisfaction pull against each other, so pick a number you can live with and stop pushing.&lt;/p&gt;

&lt;p&gt;That conclusion is usually wrong. What most teams are looking at isn't a tradeoff curve. It's two different curves that happen to overlap on the left side of the chart, and almost everyone is standing on the worse one without knowing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two curves, not one
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The capability curve.&lt;/strong&gt; Resolution goes up because the AI genuinely got better at handling more intent types. CSAT rises with it — customers get correct answers instantly instead of waiting four hours for a human to say the same thing. On this curve, satisfaction tends to peak somewhere in the &lt;strong&gt;60–80% resolution&lt;/strong&gt; range and then declines gently, because the last 20% of tickets are genuinely the ones that need judgment.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The deflection curve.&lt;/strong&gt; Resolution goes up because the path to a human got harder to find. The number on the dashboard rises identically. CSAT doesn't decline gently — it collapses.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs7u8fvajwyt0wsszcj1p.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs7u8fvajwyt0wsszcj1p.png" alt=" " width="800" height="501"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The vertical gap between those two curves is the deflection penalty, and it's the thing people mistake for a law of physics.&lt;/p&gt;

&lt;p&gt;Here's the part that makes this hard to see from the inside: CSAT is a hill, not a ramp. It's low at both ends. Under-automate and people wait too long. Over-automate and people get trapped. If you only know your CSAT is falling, you can't tell which end you're on from the metric alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the deflection penalty comes from
&lt;/h2&gt;

&lt;p&gt;Six failure modes, and most deployments have several running at once:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Suppressed escalation.&lt;/strong&gt; The handoff exists but is buried, rate-limited, or gated behind three "are you sure?" prompts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Overreach.&lt;/strong&gt; The agent attempts regulated, high-stakes, or genuinely ambiguous issues it has no business attempting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Confidently wrong answers.&lt;/strong&gt; The ticket closes. The problem doesn't.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;False resolutions.&lt;/strong&gt; A customer gives up and closes the tab. The system logs a win.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Effort and looping.&lt;/strong&gt; Three rephrasings of the same question before anything works. Customers rate effort, not accuracy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lost empathy on edge cases.&lt;/strong&gt; The people with the worst experiences are disproportionately the ones who fill out surveys.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Deflection is not resolution
&lt;/h2&gt;

&lt;p&gt;This is the crux. &lt;strong&gt;Deflection counts any conversation that didn't reach a human — including the ones where the customer rage-quit.&lt;/strong&gt; Genuine resolution excludes escalations and verifies the ticket actually stayed closed.&lt;/p&gt;

&lt;p&gt;For the same deployment, those two numbers can differ by 20–30 points. So when a vendor quotes you a resolution rate, you're not looking at a performance figure until you know which definition produced it. Half the "AI CSAT problem" discourse is people comparing a deflection number to a satisfaction number and drawing a causal arrow between them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three design choices that break the tradeoff
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Confidence-gate the automation.&lt;/strong&gt; Auto-resolve only high-confidence intents. Route ambiguity to a human by default. The instinct is to push the confidence threshold down to lift the resolution number — that's the exact move that slides you onto the deflection curve.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make escalation instant and graceful.&lt;/strong&gt; Treat a handoff as a feature, not a failure. Pass full context so the customer never repeats themselves. A fast, clean escalation produces a &lt;em&gt;higher&lt;/em&gt; CSAT than a mediocre AI answer, which means the honest version of your funnel is also the better-performing one.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Measure genuine resolution.&lt;/strong&gt; Exclude escalations. Verify the ticket stayed closed inside a defined window (72 hours is a reasonable default). Measure CSAT specifically on AI-handled tickets, not blended across every channel — blending is how a struggling AI hides behind a good phone team.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five questions for any vendor
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;Does your "resolution" metric include escalations or workflow handoffs?&lt;/li&gt;
&lt;li&gt;Are re-opens within 72 hours counted as resolution failures?&lt;/li&gt;
&lt;li&gt;Is CSAT measured on AI-handled tickets, or blended across all channels?&lt;/li&gt;
&lt;li&gt;Is resolution defined consistently across channels?&lt;/li&gt;
&lt;li&gt;Can escalation rate and resolution rate appear on the same report?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last one is the tell. If the two numbers can't sit next to each other, someone has decided you shouldn't see them together.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Before you accept that automation costs you satisfaction, check which curve you're on. Recompute resolution with escalations excluded and re-opens counted as failures, then look at CSAT on AI-handled tickets only. If genuine resolution comes in well below the deflection figure, you don't have a tradeoff to manage. You have a measurement problem to fix — and fixing it usually moves both numbers in the right direction at once.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cx</category>
    </item>
    <item>
      <title>I compared what 18 AI support vendors actually charge per resolved conversation. The spread is 25 .</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Wed, 19 Aug 2026 17:37:01 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/i-compared-what-18-ai-support-vendors-actually-charge-per-resolved-conversation-the-spread-is-25--1p1h</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/i-compared-what-18-ai-support-vendors-actually-charge-per-resolved-conversation-the-spread-is-25--1p1h</guid>
      <description>&lt;h1&gt;
  
  
  I compared what 18 AI support vendors actually charge per resolved conversation. The spread is 25×.
&lt;/h1&gt;

&lt;p&gt;&lt;em&gt;Disclosure: I work at Aissist.io, which is one of the 18 vendors in this benchmark. All figures come from public pricing pages (verified July 2026) or third-party procurement data. Full methodology and the complete table are in the &lt;a href="https://aissist.io/industries/ai-agent-pricing-benchmark-2026" rel="noopener noreferrer"&gt;original writeup&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Headline prices for AI support agents cluster politely between $0.50 and $2.00. That number is close to useless. The only figure that matters is &lt;strong&gt;effective cost per resolved conversation&lt;/strong&gt; — total spend divided by conversations actually resolved, including seats and platform fees.&lt;/p&gt;

&lt;p&gt;Once you normalize for that, the range runs from about $0.13 to $3.33+.&lt;/p&gt;

&lt;h2&gt;
  
  
  The pricing model matters more than the rate
&lt;/h2&gt;

&lt;p&gt;Five models are in play:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Per resolution&lt;/strong&gt; — you pay when the AI finishes the job. Everything hinges on the definition of "resolved."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per conversation&lt;/strong&gt; — you pay whether or not it worked. The vendor gets paid for failures.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per session&lt;/strong&gt; — one real issue often spans 3–4 billable sessions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flat per ticket&lt;/strong&gt; — cheap, but it's an add-on layer; your helpdesk bill continues underneath.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enterprise custom&lt;/strong&gt; — platform fee plus negotiated usage, typically $50K–$600K+/year.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At a 60% resolution rate, a $2.00 per-conversation fee is $3.33 per actual resolution. The headline lied by 67%.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 10,000 monthly conversations at 70% resolution actually costs
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Vendor&lt;/th&gt;
&lt;th&gt;Est. monthly&lt;/th&gt;
&lt;th&gt;Effective $/resolution&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;My AskAI&lt;/td&gt;
&lt;td&gt;~$1,000 + helpdesk&lt;/td&gt;
&lt;td&gt;~$0.14 + helpdesk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;eesel AI&lt;/td&gt;
&lt;td&gt;~$4,000 + helpdesk&lt;/td&gt;
&lt;td&gt;~$0.57 + helpdesk&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aissist.io&lt;/td&gt;
&lt;td&gt;~$4,200&lt;/td&gt;
&lt;td&gt;~$0.60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HubSpot Breeze&lt;/td&gt;
&lt;td&gt;~$4,400&lt;/td&gt;
&lt;td&gt;~$0.63&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Intercom Fin&lt;/td&gt;
&lt;td&gt;~$7,780&lt;/td&gt;
&lt;td&gt;~$1.11&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gorgias&lt;/td&gt;
&lt;td&gt;~$10,000+&lt;/td&gt;
&lt;td&gt;~$1.43&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Decagon&lt;/td&gt;
&lt;td&gt;~$14,070&lt;/td&gt;
&lt;td&gt;~$2.01&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zendesk AI&lt;/td&gt;
&lt;td&gt;~$14,500&lt;/td&gt;
&lt;td&gt;~$2.07&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sierra&lt;/td&gt;
&lt;td&gt;~$15,000–20,000&lt;/td&gt;
&lt;td&gt;~$2.14–2.86&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Salesforce Agentforce&lt;/td&gt;
&lt;td&gt;~$22,000+&lt;/td&gt;
&lt;td&gt;~$3.14+&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same 7,000 resolved conversations. Over $200,000/year of difference.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fine print that moves the number
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Assumed resolutions&lt;/strong&gt; — Intercom Fin bills when a customer stops replying, which includes the ones who gave up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Double billing&lt;/strong&gt; — Gorgias charges an AI-resolved conversation as both a ticket and a resolution unless it's escalated within 72 hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session inflation&lt;/strong&gt; — Freddy's per-session rate looks like $0.10–$0.49, but published resolution rates of 23–30% mean 3–4 sessions per real fix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Uncapped overages&lt;/strong&gt; — Zendesk has auto-billed everything above committed volume at full rate since January 2026, no cap, no warning.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Six questions to ask any vendor
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;What exactly triggers a charge? Get it in writing.&lt;/li&gt;
&lt;li&gt;What happens when the AI fails — do you still pay?&lt;/li&gt;
&lt;li&gt;Is there a floor (minimums, annual commitments)?&lt;/li&gt;
&lt;li&gt;Is there a ceiling on overages?&lt;/li&gt;
&lt;li&gt;What's the all-in stack: seats, platform, onboarding, helpdesk underneath?&lt;/li&gt;
&lt;li&gt;Who audits the resolution rate? If the vendor grades its own homework, ask for independent reporting.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For context: a human-handled ticket runs $6–$12. Most of these are still a bargain. But normalize every quote to cost per resolution first — then ask who's counting.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Nothing here is a quote or a contractual commitment. Verify with the vendor before signing.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>A benchmark tool for AI customer service — would love this community to poke holes in it</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Sun, 12 Jul 2026 18:01:15 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/a-benchmark-tool-for-ai-customer-service-would-love-this-community-to-poke-holes-in-it-5gm6</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/a-benchmark-tool-for-ai-customer-service-would-love-this-community-to-poke-holes-in-it-5gm6</guid>
      <description>&lt;p&gt;How is your AI support actually doing?&lt;/p&gt;

&lt;p&gt;Not the number in the vendor's slide deck. Not the demo. Yours — on your real traffic, this quarter.&lt;/p&gt;

&lt;p&gt;If you run AI customer service, you've probably tried to answer some version of these four questions:&lt;/p&gt;

&lt;p&gt;→ Is it truly resolving tickets, or just deflecting them until the customer gives up or a human quietly cleans it up?&lt;br&gt;
→ What's the ROI really looking like once you net out the cost of the platform, the setup, and the escalations?&lt;br&gt;
→ Is there still room to improve — or are you already near the ceiling for the kind of traffic you handle?&lt;br&gt;
→ And how do you actually compare with everyone else running the same play in your industry?&lt;/p&gt;

&lt;p&gt;Here's the uncomfortable part: almost nobody can answer these with confidence. And it's not because support leaders aren't paying attention. It's because the ground truth doesn't exist in any shared, honest form.&lt;/p&gt;

&lt;p&gt;Start with the word "resolution." It sounds precise. It isn't. One vendor counts a ticket as resolved the moment the bot sends a reply. Another counts it only if the customer never writes back. A third quietly folds deflection — the customer bouncing off a help article — into the same bucket and calls it automation. So when you see "85% resolution" on a landing page, you have no idea whether that's a genuinely closed issue or a cleverly drawn chart. The definition bends to flatter whoever is publishing it.&lt;/p&gt;

&lt;p&gt;That makes comparison almost impossible. If you're sitting at 60%, is that world-class for your vertical, or embarrassing? You genuinely can't tell — because there's no shared denominator, no neutral reference point, and every benchmark you can find was published by someone with a reason to make their own number look good.&lt;/p&gt;

&lt;p&gt;That lack of transparency is the problem we set out to close.&lt;/p&gt;

&lt;p&gt;We built a benchmark tool, and we deliberately backed it with two different kinds of data:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Public data — independent 2026 cross-program aggregates, pulled from sources that aren't trying to sell you anything.&lt;/li&gt;
&lt;li&gt;Private data — real field deployments, where we can see what actually happens once a system is live on messy, real-world traffic.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The reason we use both matters. Public data keeps us honest and broad. Private data keeps us grounded in what really happens after launch — not the pilot, not the demo, but month six, when the easy intents are automated and the hard ones are all that's left.&lt;/p&gt;

&lt;p&gt;Here's how it works. You pick your industry — ecommerce, fintech, SaaS, travel, healthcare, telecom, and so on — and the tool shows you the realistic resolution and CSAT range for that vertical. Not a single vanity figure, but a band: the low end, the median, and the top-10% world-class mark. Then you can enter your own current resolution rate and CSAT, and it places you directly on that scale. In about ten seconds you can see whether you're lagging, sitting at the median, or genuinely at the front of your industry.&lt;/p&gt;

&lt;p&gt;And because a number without context is just anxiety, the tool also breaks down the four factors that decide where any given deployment actually lands:&lt;/p&gt;

&lt;p&gt;→ Industry complexity. How ambiguous, regulated, emotional, or multi-step your intents are. This sets the ceiling. Structured, data-rich intents like order status resolve high; regulated or emotional ones resolve far lower, no matter which vendor you use.&lt;/p&gt;

&lt;p&gt;→ System capability. How strong the underlying platform is — multi-agent reasoning, backend actions, retrieval, guardrails. A more capable system reaches higher within the same industry band and holds quality as complexity rises.&lt;/p&gt;

&lt;p&gt;→ Maturity of your assets and playbooks. The state of your knowledge base, SOPs, and escalation rules. Most deployments launch around 40–50% and climb past 60% over six to twelve months as those assets mature. If you're new, low isn't failure — it's the starting line.&lt;/p&gt;

&lt;p&gt;→ Transparency of your data. How accessible the orders, accounts, and history are that the AI needs to actually resolve a ticket. When the system can see everything, it resolves. When data is siloed, even simple intents stall.&lt;/p&gt;

&lt;p&gt;The honest caveat: these ranges are directional, not audited. Definitions of "resolution" vary by source, so we treat the numbers as a realistic target band — something to steer by — rather than a certified figure. We'd rather tell you that plainly than pretend to a precision the whole industry can't actually deliver yet.&lt;/p&gt;

&lt;p&gt;This is a living benchmark. We'll keep updating it as more public data lands and more deployments mature, and we expect the ranges to sharpen over time. If your team has been looking for a bit of honest transparency in a space that badly lacks it, we hope this gives you a real reference point to plan against — whether or not you ever talk to us.&lt;/p&gt;

&lt;p&gt;Try it here: &lt;a href="https://aissist.io/tools/benchmark" rel="noopener noreferrer"&gt;https://aissist.io/tools/benchmark&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94xi1s4uunzx3vitdf9v.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F94xi1s4uunzx3vitdf9v.png" alt=" " width="799" height="377"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;And I'd genuinely value your feedback. If the range for your industry feels off compared to what you're seeing in your own numbers, tell me — that's exactly the kind of signal that makes the next version better for everyone.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The fix for bad AI isn't a bigger brain. It's better structure.</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Thu, 09 Jul 2026 23:30:25 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/the-fix-for-bad-ai-isnt-a-bigger-brain-its-better-structure-3a0c</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/the-fix-for-bad-ai-isnt-a-bigger-brain-its-better-structure-3a0c</guid>
      <description>&lt;p&gt;Most AI support bots top out at 30-50% resolution and get stuck around 3.5/5.0 CSAT.&lt;/p&gt;

&lt;p&gt;Customers describe them with one word: robotic.&lt;/p&gt;

&lt;p&gt;We kept hitting the same wall, and the reason turned out to be structural. A single generic model tries to force every refund, every connectivity issue, every warranty claim, every shipping question through one path. But real business problems aren't linear. They're full of ambiguity, overlapping systems, and incomplete information. Rigid decision trees and one big model both break the moment real complexity shows up.&lt;/p&gt;

&lt;p&gt;So we stopped trying to build one model that knows everything.&lt;/p&gt;

&lt;p&gt;Instead: divide and conquer. A sub-agent is a specialist focused on one domain — refunds, connectivity, shipping, warranty, finance, product damage. On each execution, a super agent plans the work, activates the right specialists, gathers facts from them, and makes the final call. Multiple sub-agents can contribute to the same request and cross-check each other — which is exactly why the output is more reliable than any single-agent system.&lt;/p&gt;

&lt;p&gt;The results across Q1 2026 deployments:&lt;/p&gt;

&lt;p&gt;83% resolution (vs. the 30-50% ceiling)&lt;br&gt;
4.8/5.0 CSAT (vs. ~3.5)&lt;/p&gt;

&lt;p&gt;And because each sub-agent is measurable — traffic, resolution, NPS, CSAT at the specialist level — you don't just learn whether the AI works. You learn about your own products, customers, and processes.&lt;/p&gt;

&lt;p&gt;The counterintuitive part: running multiple agents per request should cost more than a single model. After continuous engine optimization, it's often more cost-effective than single-model alternatives — while doing deeper reasoning.&lt;/p&gt;

&lt;p&gt;The lesson we keep relearning: reliability in AI doesn't come from a bigger model. It comes from better structure.&lt;/p&gt;

&lt;p&gt;How the architecture works: &lt;a href="https://aissist.io/technology/multi-agent-platform" rel="noopener noreferrer"&gt;https://aissist.io/technology/multi-agent-platform&lt;/a&gt;&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>The "Moore's Law for AI" everyone promised you? It didn't happen.</title>
      <dc:creator>Mike Wei</dc:creator>
      <pubDate>Thu, 09 Jul 2026 23:24:39 +0000</pubDate>
      <link>https://dev.to/michael_wei_d93a005ecc379/the-moores-law-for-ai-everyone-promised-you-it-didnt-happen-25e4</link>
      <guid>https://dev.to/michael_wei_d93a005ecc379/the-moores-law-for-ai-everyone-promised-you-it-didnt-happen-25e4</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5sv9rs4tm9fib0x05f5y.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5sv9rs4tm9fib0x05f5y.png" alt=" " width="800" height="569"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Everyone building AI agents is about to learn the same expensive lesson: capability and economics are two different problems.&lt;/p&gt;

&lt;p&gt;I've spent 15+ years building large-scale ML systems, and here's the uncomfortable math behind agentic AI in 2026:&lt;/p&gt;

&lt;p&gt;→ Token prices haven't fallen in a straight line. The cheapest models have flattened out, and flagship pricing turned upward again this year (GPT-5.4 launched at 2x GPT-5's input price).&lt;/p&gt;

&lt;p&gt;→ Reasoning tokens are billed as output tokens — deeper thinking raises cost even when the visible answer stays short.&lt;/p&gt;

&lt;p&gt;→ One agentic request expands into planning, retrieval, verification, and tool calls. More useful work, way more tokens.&lt;/p&gt;

&lt;p&gt;Betting your unit economics on "models will get cheaper" is not a strategy. Power constraints alone make that assumption shaky — AI demand is now colliding with the electricity grid, not just GPU supply.&lt;/p&gt;

&lt;p&gt;What actually works is treating token efficiency as an architecture problem:&lt;/p&gt;

&lt;p&gt;System design is the strongest lever. A system tightly coupled to the problem spends less because it already knows the workflow, the boundaries, and the next step. Generic agents pay a "rediscovery tax" on every single run.&lt;/p&gt;

&lt;p&gt;Task optimization comes second. Routing, classification, extraction — these don't need frontier reasoning. The first big savings usually come from stopping the model from doing work it never needed to do.&lt;/p&gt;

&lt;p&gt;Custom models come last, and only at scale. Trade a recurring inference bill for a predictable training cost — but only when the task is stable and you have real evaluation discipline.&lt;/p&gt;

&lt;p&gt;This is why we treat cost as a core design constraint at Aissist, not a cleanup task for later. AI that can't scale economically fails the business case no matter how smart it is.&lt;/p&gt;

&lt;p&gt;Full analysis: &lt;a href="https://aissist.io/technology/token-efficiency" rel="noopener noreferrer"&gt;https://aissist.io/technology/token-efficiency&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
