<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Olamide Ashiru</title>
    <description>The latest articles on DEV Community by Olamide Ashiru (@alpha-dev-001).</description>
    <link>https://dev.to/alpha-dev-001</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4015682%2F5e5fce86-c695-4544-a21c-f652238ae833.png</url>
      <title>DEV Community: Olamide Ashiru</title>
      <link>https://dev.to/alpha-dev-001</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alpha-dev-001"/>
    <language>en</language>
    <item>
      <title>I Gave Qwen a Store to Run — Then Took Away Its Ability to Pick a Number</title>
      <dc:creator>Olamide Ashiru</dc:creator>
      <pubDate>Tue, 07 Jul 2026 06:31:17 +0000</pubDate>
      <link>https://dev.to/alpha-dev-001/elevate-making-qwen-the-brain-of-a-store-that-runs-itself-582p</link>
      <guid>https://dev.to/alpha-dev-001/elevate-making-qwen-the-brain-of-a-store-that-runs-itself-582p</guid>
      <description>&lt;p&gt;&lt;em&gt;An autonomous commerce agent where the merchant just approves. Built for the Global AI Hackathon with Qwen Cloud — *Track 4: Autopilot Agent.&lt;/em&gt;**&lt;/p&gt;




&lt;p&gt;I ran one discount decision through Qwen across seven scenarios — different costs, different margins, different price ceilings, genuinely different economics in each.&lt;/p&gt;

&lt;p&gt;It said &lt;strong&gt;10%&lt;/strong&gt;. Every single time.&lt;/p&gt;

&lt;p&gt;Not wrong in a dramatic way — 10% is a fine, safe-sounding number. That's exactly the problem. A confidently generic answer, applied to seven situations that each needed a different one, is the kind of thing that quietly bleeds a real merchant while looking completely reasonable on screen.&lt;/p&gt;

&lt;p&gt;I only &lt;em&gt;know&lt;/em&gt; it said 10% every time because the version of the agent with brakes drove those same seven scenarios to a spread — &lt;strong&gt;2.89%, 40%, 11.11%, 10%, 10%, 15%, and one action with no discount to clamp at all&lt;/strong&gt; — a range from 2.89% to 40%, because the economics genuinely differed. Same model, same facts, run once through a live harness. The only difference was whether it was allowed to pick the number.&lt;/p&gt;

&lt;p&gt;That gap — between what the model &lt;em&gt;proposes&lt;/em&gt; and what it's &lt;em&gt;allowed to do&lt;/em&gt; — is the whole thing I built. Elevate is an AI that designs a storefront from a logo, stocks and prices it from a folder of photos, and then &lt;em&gt;runs&lt;/em&gt; it from live shopper behavior, with the merchant as the human-in-the-loop. But the lesson underneath it is smaller and more reusable than "AI runs a store":&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Let the model decide &lt;em&gt;what&lt;/em&gt; to do. Never let it invent the &lt;em&gt;number&lt;/em&gt;. Put your guardrail on the number.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I ended up applying that rule three separate times, and every time it was the difference between a demo and something I'd hand real money to.&lt;/p&gt;




&lt;h3&gt;
  
  
  TL;DR
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What it is:&lt;/strong&gt; an autopilot commerce agent — Qwen builds the store, stocks it, prices it, watches it, and proposes actions; the merchant approves the exceptions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What's novel:&lt;/strong&gt; (1) a role-scoped &lt;strong&gt;swarm&lt;/strong&gt; that proposes and a merchant who disposes; (2) a &lt;strong&gt;structural guard + 3-layer interceptor&lt;/strong&gt; the agent literally cannot bypass; (3) a closed &lt;strong&gt;learning loop&lt;/strong&gt; — every decision reads back real revenue, so it gets smarter &lt;em&gt;per store&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The principle:&lt;/strong&gt; the model decides &lt;em&gt;what&lt;/em&gt;; deterministic code owns every &lt;em&gt;number&lt;/em&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stack:&lt;/strong&gt; Qwen (&lt;code&gt;qwen-vl-max&lt;/code&gt; + &lt;code&gt;qwen-max&lt;/code&gt;) on Alibaba Cloud ECS + OSS. Two models, seven jobs, one learning loop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Repo:&lt;/strong&gt; &lt;a href="https://github.com/Alpha-dev-001/elevate-hackathon" rel="noopener noreferrer"&gt;https://github.com/Alpha-dev-001/elevate-hackathon&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  The principle, three times
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Pricing.&lt;/strong&gt; Resale inventory has no honest MSRP — web-search "the usual price" and you get $400 designer retail for a slide a local reseller sells for a fraction of that. A confidently-wrong price is worse than no price. So Qwen never picks the number: it anchors to a merchant-set baseline and nudges by the premium-ness it can actually &lt;em&gt;see&lt;/em&gt; in the photo, clamped so a hallucinated figure can't reach the storefront. It decides &lt;em&gt;how premium this looks&lt;/em&gt;. It does not decide &lt;em&gt;what it costs&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Revenue estimates.&lt;/strong&gt; Left to invent "how much will this promo make," Qwen answered the same trigger with $760, then $175, then $285 — confident and meaningless. So I took the number away. The estimate is now &lt;em&gt;computed&lt;/em&gt; from real signals: the size of the anomaly it actually detected × the store's real average price × a tunable rate. Qwen decides &lt;em&gt;there's an opportunity here&lt;/em&gt;. Arithmetic decides &lt;em&gt;how big&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Discounts.&lt;/strong&gt; This is where the brakes get physical. Before an action is even considered, a &lt;strong&gt;Layer-0 structural guard&lt;/strong&gt; makes an illegal one &lt;em&gt;unrepresentable&lt;/em&gt; — a negative discount, an over-100% discount, a zero price, a phantom product simply can't be constructed into an action. Whatever survives hits a three-layer interceptor neither the merchant &lt;em&gt;nor Qwen&lt;/em&gt; can bypass:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Layer 1 — Brand Guard:&lt;/strong&gt; Qwen-authored rules, checked client-side with zero latency, so a clashing color warns the instant you pick it — in Qwen's own words.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer 2 — Business Constraints:&lt;/strong&gt; margin floors and discount ceilings, auto-clamped with a visible reason. In testing, a decision cycle proposed a &lt;strong&gt;75%&lt;/strong&gt; win-back; it was clamped to that store's &lt;strong&gt;40%&lt;/strong&gt; ceiling before it could reach a customer. Qwen's number, the merchant's limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Layer 3 — System Safety:&lt;/strong&gt; anything that would sell below unit cost is hard-blocked. No clamp, no exception — refused.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The agent can propose freely &lt;em&gt;precisely because&lt;/em&gt; the brakes are real and it can't touch them. That's what makes autopilot safe to hand real control.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;[screenshot: the Layer-2 clamp — Qwen's 75% struck through, clamped to 40%, reason shown]&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What it actually does (the fast version)
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;A logo becomes a store.&lt;/strong&gt; One upload. &lt;code&gt;qwen-vl-max&lt;/code&gt; reads it into structured brand DNA — colors, mood, industry cues — and &lt;code&gt;qwen-max&lt;/code&gt; turns that into a full brand plus a &lt;strong&gt;LayoutDSL&lt;/strong&gt;: a JSON description of the store's &lt;em&gt;structure&lt;/em&gt;, not a template with the colors swapped. Forty logos give forty visually distinct storefronts. And because a model returning structure &lt;em&gt;will&lt;/em&gt; hallucinate an invalid variant or eight sections where there should be two, the renderer never trusts raw output: it passes through type-coercion, structural normalization, and a deterministic &lt;strong&gt;brand-seeded fallback&lt;/strong&gt; so that even a total Qwen failure yields a distinct, on-brand store. A broken storefront is impossible by construction.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;[screenshot: three generated storefronts side by side — Burger Blitz (light, editorial), Xair (dark, structured), Owoyemi of Offa — same pipeline, no shared template]&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A folder of photos becomes a catalog.&lt;/strong&gt; Inventory doesn't arrive as a tidy spreadsheet — it arrives as 98 phone photos of branded footwear, names on nothing, prices nowhere. (I know because a real seller handed me exactly that.) &lt;code&gt;qwen-vl-max&lt;/code&gt; reads every photo into a sellable name, colorways, brand-voice description, category, and a &lt;em&gt;suggested&lt;/em&gt; price. It also stays honest about its own eyes: it once misread a "SUICOKE" strap as "Suicide," fully confident — which is exactly why low-confidence reads are drafted &lt;strong&gt;inactive&lt;/strong&gt; and handed back for a human glance instead of going live. The agent does the 98× of tedious work; the human rules on the handful that matter.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Shopper behavior becomes action.&lt;/strong&gt; Once live, events stream over WebSocket into Redis. A deterministic threshold watches for patterns — a cart-abandon surge, a velocity spike — and Qwen also proactively checks pricing, scarcity, and catalog health on a schedule. When something crosses, it doesn't hand the whole problem to one do-everything prompt. It routes to one of &lt;strong&gt;four role-scoped specialists&lt;/strong&gt; — a Pricing Strategist, a Sales Rep, a Store Curator, an Inventory Overseer — each holding only its own tools, each able to &lt;em&gt;escalate&lt;/em&gt; to another when the real fix is outside its lane. That specialist runs a decision cycle, and &lt;strong&gt;before it decides, it reads its own memory&lt;/strong&gt;: every past action for &lt;em&gt;this&lt;/em&gt; store and whether it drove revenue. It even carries a quantified learned stance — "kept offers averaged 9% off, dismissed ones 35%, so lead nearer 9%" — so its proposals converge on what this specific merchant actually approves.&lt;/p&gt;

&lt;p&gt;It returns one &lt;strong&gt;option card&lt;/strong&gt;: the play, the trigger, the (computed) estimated revenue, a confidence score, a brand-safety check. Then it stops. The merchant taps &lt;strong&gt;Approve&lt;/strong&gt;, and the storefront morphs, live. Qwen proposes; the merchant disposes; nothing happens behind the owner's back.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;[screenshot: an option card in the merchant terminal — play, trigger, estimated revenue, confidence, brand check]&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;And the loop closes.&lt;/strong&gt; When a promo resolves, a background observer counts the orders attributed to it and writes a memory entry the next decision reads. In a live run on the demo store, a cart-abandon surge → an approved win-back → a shopper checking out under it attributed &lt;strong&gt;$74.70&lt;/strong&gt; to that one decision — revenue moving $150.50 → $225.20, the platform's 10% fee ($7.47) computed against it — measured on the exact path a merchant would watch, not asserted by the model.&lt;/p&gt;




&lt;h2&gt;
  
  
  The bugs all lived in the seams
&lt;/h2&gt;

&lt;p&gt;Every hard bug had the same shape: each component was individually correct, individually tested — and &lt;em&gt;nothing connected them&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;The one that still bothers me: our flagship dynamic-pricing feature — the one that reprices a product from its own sales history, earns trust, and eventually auto-applies — had &lt;strong&gt;never once executed in the live system.&lt;/strong&gt; The reasoning, the reversion logic, the trust graduation: all written, all unit-tested, all green. But the daily job that rolls raw behavior into the price-history table was never wired into any background loop. So that table sat at zero rows forever, and the eligibility check that reads it was therefore &lt;em&gt;always false&lt;/em&gt;. An entire headline feature, dead behind a passing test suite — surfaced only by querying the live database and asking "why are there no history rows?"&lt;/p&gt;

&lt;p&gt;Its twin: the attribution dashboard — the closing beat of the whole pitch, &lt;em&gt;"this action drove $X"&lt;/em&gt; — was quietly reading &lt;strong&gt;$0&lt;/strong&gt;, because it joined orders on a promo's human-readable &lt;em&gt;label&lt;/em&gt; instead of its &lt;em&gt;id&lt;/em&gt;. That silent mismatch didn't just zero the number; it fed the memory loop false "no conversions," teaching Qwen its &lt;em&gt;working&lt;/em&gt; actions had failed.&lt;/p&gt;

&lt;p&gt;Same lesson twice: &lt;strong&gt;the happy path is not the interesting path.&lt;/strong&gt; A feature isn't done when its logic passes tests — it's done when something on the live path actually &lt;em&gt;calls&lt;/em&gt; it. The most dangerous failures don't throw. They quietly report zero. (I added a live end-to-end test that seeds a real purchase and proves the feature flips on — because "unit-tested" is exactly what lied to me.)&lt;/p&gt;




&lt;h2&gt;
  
  
  Does it pencil out?
&lt;/h2&gt;

&lt;p&gt;Both were design constraints, not afterthoughts. Building a full store — logo read, brand, icons, ~90-product catalog — costs on the order of &lt;strong&gt;a few cents&lt;/strong&gt; in tokens, once, then it's cached forever. After launch Qwen doesn't busy-poll; it fires on a real anomaly or a bounded scheduled check, and sends a telemetry &lt;em&gt;diff&lt;/em&gt;, not the whole state. A decision cycle lands in seconds. Cost scales with the events that matter, not the hours the store is open — the only way "an agent per merchant" is economically sane.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Deployment:&lt;/strong&gt; FastAPI, Postgres, and Redis as Docker containers on a single Alibaba Cloud ECS instance, with Alibaba OSS for logo storage and Qwen Cloud for every model call.&lt;/p&gt;




&lt;h2&gt;
  
  
  The feature I built, then deliberately took away from the model
&lt;/h2&gt;

&lt;p&gt;I built graduated autonomy first: a trust streak where, once a product's small, already-safe price moves had been approved enough times, the model could apply the next one &lt;em&gt;without a tap&lt;/em&gt;. It felt like the natural endpoint of a learning agent — earn your way to acting alone.&lt;/p&gt;

&lt;p&gt;Two things killed it. First, at final review I found the streak was structurally dead — built, tested, and never actually firing, the &lt;em&gt;second&lt;/em&gt; feature I caught passing its checks while doing nothing (the seams strike again). But fixing the bug is what made me question the whole idea. Because the real problem wasn't the bug. It was the premise: &lt;strong&gt;why is the model the thing that earns the right to act unattended?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That's the same mistake as letting it pick the number, one level up. An agent grading its own trustworthiness and then promoting itself is exactly the blind spot you don't want — the human's rubber-stamp quietly becomes consent they never actively gave.&lt;/p&gt;

&lt;p&gt;So autonomy is now &lt;strong&gt;opt-in, per product.&lt;/strong&gt; The model never grants itself authority. The merchant flips a switch on a specific product — "you handle this one" — and hands over a scoped sample, revocably, when &lt;em&gt;they&lt;/em&gt; decide Elevate has earned it. The model still learns, still proposes, still gets sharper per store. It just never promotes itself. That's the actual goal of the whole system: not an AI that seizes control by clearing a threshold, but one a merchant can safely delegate to, one product at a time, so they can go run the parts of the business that need a human.&lt;/p&gt;

&lt;p&gt;Which turns out to be the same principle as everything else here, applied to power instead of price: &lt;strong&gt;let the model decide what to do. Never let it decide the number — or how much authority it has.&lt;/strong&gt; Everything trustworthy in Elevate came from that one line.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Code: &lt;a href="https://github.com/Alpha-dev-001/elevate-hackathon" rel="noopener noreferrer"&gt;https://github.com/Alpha-dev-001/elevate-hackathon&lt;/a&gt; — MIT licensed. Built for the Global AI Hackathon Series with Qwen Cloud, Track 4: Autopilot Agent.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>python</category>
    </item>
  </channel>
</rss>
