<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Toheeb Olanrewaju Olagoke</title>
    <description>The latest articles on DEV Community by Toheeb Olanrewaju Olagoke (@olacode).</description>
    <link>https://dev.to/olacode</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4071341%2F75bb00aa-e115-4643-a902-f5bd13744dfc.png</url>
      <title>DEV Community: Toheeb Olanrewaju Olagoke</title>
      <link>https://dev.to/olacode</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/olacode"/>
    <language>en</language>
    <item>
      <title>Building Co-Shop: A Shared Cart for Humans and AI Agents with WebMCP</title>
      <dc:creator>Toheeb Olanrewaju Olagoke</dc:creator>
      <pubDate>Wed, 09 Sep 2026 21:43:01 +0000</pubDate>
      <link>https://dev.to/olacode/building-co-shop-a-shared-cart-for-humans-and-ai-agents-with-webmcp-5p8</link>
      <guid>https://dev.to/olacode/building-co-shop-a-shared-cart-for-humans-and-ai-agents-with-webmcp-5p8</guid>
      <description>&lt;p&gt;&lt;strong&gt;The problem with "agentic" web apps today&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask most AI agents to shop for you online right now, and here's roughly what happens under the hood: the agent opens a headless browser, loads the page, tries to parse the DOM, guesses which  is the "Add to Cart" one, clicks it, waits, re-parses, and hopes nothing changed in the layout since the model was trained on how that site "usually" looks.&lt;/p&gt;


&lt;p&gt;It works sometimes. It's also slow, brittle, and completely opaque to the human sitting there. You ask an agent to "add a few vegetarian dinners to my cart," it goes away for a minute, and you get a summary back: "Done! I added 3 items." You had no idea what it was doing while it did it, and if it clicked the wrong thing, you find out after the fact.&lt;/p&gt;

&lt;p&gt;WebMCP is a proposed web standard that removes the guesswork entirely. Instead of an agent reverse engineering your UI, your page tells the agent what it can do:&lt;/p&gt;

&lt;p&gt;document.modelContext.registerTool({&lt;/p&gt;

&lt;p&gt;name: "add_to_cart",&lt;/p&gt;

&lt;p&gt;description: "Add a product to the cart",&lt;/p&gt;

&lt;p&gt;inputSchema: {&lt;/p&gt;

&lt;pre class="highlight plaintext"&gt;&lt;code&gt;type: "object",

properties: {

  productId: { type: "string" },

  quantity: { type: "number" },

},

required: ["productId"],
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;},&lt;/p&gt;

&lt;p&gt;async execute({ productId, quantity }) {&lt;/p&gt;

&lt;pre class="highlight plaintext"&gt;&lt;code&gt;// your actual cart logic
&lt;/code&gt;&lt;/pre&gt;

&lt;p&gt;},&lt;/p&gt;

&lt;p&gt;});&lt;/p&gt;

&lt;p&gt;An agent visiting your page right now, in ChatGPT's WebMCP-enabled desktop browser, or in Chrome with the #enable-webmcp-testing flag can call add_to_cart directly. No clicking, no scraping, no guessing.&lt;/p&gt;

&lt;p&gt;I wanted to see what this actually unlocks beyond "agent can now click buttons faster," so I built Co-Shop for OpenAI's WebMCP Challenge: a grocery storefront where a human and an AI agent share the literal same cart, live, in the same browser tab.&lt;/p&gt;

&lt;p&gt;Here's what I learned building it.&lt;br&gt;
The core idea: one state, two actors&lt;br&gt;
The easy version of this project would have been "agent has its own cart-building tool, human has a UI, sync them somehow." I didn't want that. The whole point of WebMCP, I think, is that there doesn't need to be a sync step at all if the agent's tool and the human's UI both mutate the same piece of state, they're never out of sync in the first place.&lt;/p&gt;

&lt;p&gt;So in Co-Shop, app/page.js is a single client component holding the cart in React state, and both the UI's click handlers and the WebMCP tools' execute functions call the exact same helper functions:&lt;/p&gt;

&lt;p&gt;function addToCart(productId, quantity, actor) {&lt;/p&gt;

&lt;p&gt;const product = PRODUCTS.find((p) =&amp;gt; p.id === productId);&lt;/p&gt;

&lt;p&gt;if (!product) return { ok: false, error: &lt;code&gt;No product with id "${productId}".&lt;/code&gt; };&lt;/p&gt;

&lt;p&gt;const qty = Math.max(1, Math.floor(quantity || 1));&lt;/p&gt;

&lt;p&gt;setCart((prev) =&amp;gt; {&lt;/p&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;const existing = prev.find((i) =&amp;gt; i.productId === productId);

if (existing) {

  return prev.map((i) =&amp;gt;

    i.productId === productId ? { ...i, quantity: i.quantity + qty, addedBy: actor } : i

  );

}

return [...prev, { productId, quantity: qty, addedBy: actor }];
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;});&lt;/p&gt;

&lt;p&gt;logActivity(actor, &lt;code&gt;added ${qty}× ${product.emoji} ${product.name} to the cart&lt;/code&gt;);&lt;/p&gt;

&lt;p&gt;return { ok: true, product: product.name, quantity: qty };&lt;/p&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;The only difference between a human clicking "Add" and an agent calling the add_to_cart tool is the actor string passed in: "human" vs "agent". That one parameter is what drives the purple "AGENT" badge, the pulse animation, and the activity feed entry. Everything else the state, the total, the render is identical.&lt;br&gt;
Registering the tools&lt;br&gt;
All seven tools are registered in a single useEffect, using an AbortController so cleanup is automatic if the component unmounts:&lt;/p&gt;

&lt;p&gt;_useEffect(() =&amp;gt; {&lt;/p&gt;

&lt;p&gt;if (typeof window === "undefined" || !("modelContext" in document)) {&lt;/p&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;setWebmcpStatus("unsupported");

return;
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;const controller = new AbortController();&lt;/p&gt;

&lt;p&gt;const opts = { signal: controller.signal };&lt;/p&gt;

&lt;p&gt;async function registerAll() {&lt;/p&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;await document.modelContext.registerTool(

  {

    name: "search_products",

    description: "Search the Co-Shop grocery catalog...",

    inputSchema: {

      type: "object",

      properties: {

        query: { type: "string" },

        category: { type: "string" },

        tags: { type: "array", items: { type: "string" } },

        maxPrice: { type: "number" },

      },

    },

    async execute({ query, category, tags, maxPrice } = {}) {

      const results = findProducts({ query, category, tags, maxPrice });

      return { content: [{ type: "text", text: JSON.stringify({ count: results.length, results }) }] };

    },

  },

  opts

);

// ... six more tools registered the same way

setWebmcpStatus("ready");
&lt;/code&gt;&lt;/pre&gt;
&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;registerAll();&lt;/p&gt;

&lt;p&gt;return () =&amp;gt; controller.abort(); // unregisters everything&lt;/p&gt;

&lt;p&gt;}, []);_&lt;/p&gt;

&lt;p&gt;Two details tripped me up here, and I think they're worth calling out for anyone building their first WebMCP tool:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Stale closures. Tools are registered once, on mount. If your execute function captures cart directly from a useState value at registration time, it'll keep reading that original empty cart forever. React state updates don't retroactively update already-created closures. The fix is to route every read and write through setCart (the function, not the value) and derive snapshots on demand rather than trusting a captured variable.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Return shape matters. WebMCP's execute functions are expected to return a content-array shape:&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;{ content: [{ type: "text", text: "..." }] }&lt;/p&gt;

&lt;p&gt;I originally just returned plain objects. It technically "worked" in my own manual testing (since I controlled both ends), but once a real agent called the tools, structured JSON in that text field not a bare string is what let the agent actually reason over the result (e.g., knowing exactly which 5 of 7 requested dinners fit a budget, and why).&lt;br&gt;
The tool that mattered most: compound actions&lt;br&gt;
The single tools (add_to_cart, remove_from_cart, get_cart) are the "hello world" of WebMCP pretty much the exact example in OpenAI's own challenge brief. The one that actually made the demo interesting was a compound tool:&lt;/p&gt;

&lt;p&gt;async execute({ days, dietary, avoidTags, maxBudget }) {&lt;/p&gt;

&lt;p&gt;return textResult(planDinners({ days, dietary, avoidTags, maxBudget }));&lt;/p&gt;

&lt;p&gt;}&lt;/p&gt;

&lt;p&gt;plan_dinners takes a natural-language-shaped request "7 days, vegetarian, under $40" and internally does the filtering, picks distinct matching meals, and calls addToCart for each one, all in a single tool round-trip. This is the difference between an agent doing one thing at a time forever, and an agent actually executing a plan. When I tested it against ChatGPT's in-app browser, asking it to "plan a week of vegetarian dinners under $40" resulted in exactly this response, generated by the tool itself:&lt;/p&gt;

&lt;p&gt;"Added 5 vegetarian dinners to the cart for $39.50. Seven don't fit under the $40 limit."&lt;/p&gt;

&lt;p&gt;That sentence isn't the LLM improvising it's structured output straight from plan_dinners' return value, which the model then relayed. That's the part of WebMCP I find genuinely different from prompt-and-hope agent behavior: the agent isn't guessing what happened, it's told exactly what happened, in a format it can act on.&lt;br&gt;
Testing against a real agent&lt;br&gt;
Two ways to verify a WebMCP integration actually works, in increasing order of "does this prove anything real":&lt;/p&gt;

&lt;p&gt;document.modelContext exists in the console. Proves the API surface is there. Doesn't prove your tools are being called correctly.&lt;br&gt;
A stubbed registerTool in a headless test. I did this with Playwright before ever touching a real agent stub document.modelContext.registerTool to just store the tool object, then call .execute() directly in a script. This caught real bugs (a quantity-update edge case) with instant feedback, no agent round-trip needed.&lt;br&gt;
An actual agent client. For me, that was Chrome with chrome://flags/#enable-webmcp-testing enabled (confirms registration my status pill flipped to "Agent tools live") and, more meaningfully, ChatGPT's desktop app, which has a built-in WebMCP-aware browser. Opening my deployed URL there and typing a plain-English request was the first time I saw an actual model decide, on its own, which tool to call and with what arguments.&lt;/p&gt;

&lt;p&gt;That third step is the one I'd tell anyone building a WebMCP project not to skip, even under deadline pressure. Stubbed tests prove your code is correct. They don't prove an agent will actually choose to call your plan_dinners tool instead of, say, five separate add_to_cart calls, or prove your tool descriptions are clear enough for the model to pick the right one at all.&lt;br&gt;
What I'd do differently&lt;br&gt;
If I kept building this, the next thing I'd add is a compare_products tool right now an agent can search and filter, but can't ask "which of these is the better value per serving," which feels like the natural next compound action after plan_dinners. I'd also want to test with more than one agent client side-by-side, since tool-calling behavior (how eagerly a model reaches for a compound tool vs. chaining primitives) seems to vary meaningfully between clients.&lt;/p&gt;

&lt;p&gt;Try it yourself&lt;br&gt;
Live app: co-shop-gules.vercel.app&lt;br&gt;
Source (MIT licensed): github.com/Olacode01/co-shop&lt;/p&gt;

&lt;p&gt;If you have a WebMCP-enabled browser handy, open the live link and ask an agent to "plan a week of vegetarian dinners under $40." Watching your own cart fill up live, tagged by who added what, is a much better way to understand what WebMCP changes than reading about it.&lt;/p&gt;


</description>
      <category>javascript</category>
      <category>webmcp</category>
      <category>nextjs</category>
      <category>ai</category>
    </item>
    <item>
      <title>Building VeraMove: An AI That Calls Three Movers, Catches Hidden Fees, and Negotiates a Better Deal.</title>
      <dc:creator>Toheeb Olanrewaju Olagoke</dc:creator>
      <pubDate>Mon, 10 Aug 2026 13:38:14 +0000</pubDate>
      <link>https://dev.to/olacode/building-veramove-an-ai-that-calls-three-movers-catches-hidden-fees-and-negotiates-a-better-deal-lm7</link>
      <guid>https://dev.to/olacode/building-veramove-an-ai-that-calls-three-movers-catches-hidden-fees-and-negotiates-a-better-deal-lm7</guid>
      <description>&lt;p&gt;28 million Americans move every year, in a $20B+ market made up of over 16,000 small moving companies. Real quotes for one identical 45-mile move have been documented ranging from $1,158 to $6,506, a 5.6x spread for the same job. Sight-unseen phone estimates are 40% more likely to end in a bill above the original quote. That's the problem our four-person team set out to solve at a recent Hack-Nation × ElevenLabs hackathon, and this is the story of how we built VeraMove: a voice-and-document intake pipeline that locks a single move specification, calls three vendors with it, catches the fees they don't mention up front, negotiates one quote against another, and hands back a ranked, evidence-backed recommendation. The core loop: The whole product boils down to one sequence, and we treated it as the single thing that had to work before anything else mattered:&lt;/p&gt;

&lt;p&gt;The full source is public: &lt;a href="https://github.com/zukhriddingit/VeraMove" rel="noopener noreferrer"&gt;github.com/zukhriddingit/VeraMove&lt;/a&gt;; it's a mock-first hackathon starter, no API keys or accounts required to run it yourself.&lt;/p&gt;

&lt;p&gt;Voice or document intake&lt;/p&gt;

&lt;p&gt;→ confirmed, version-locked JobSpec&lt;/p&gt;

&lt;p&gt;→ three parallel vendor calls&lt;/p&gt;

&lt;p&gt;→ itemized quotes with hidden-fee detection&lt;/p&gt;

&lt;p&gt;→ negotiation using a verified competing quote as leverage&lt;/p&gt;

&lt;p&gt;→ ranked recommendation with transcript and recording evidence. Everything else  UI polish, awards strategy, video scripts- was explicitly secondary. "A working call beats a polished interface" became something close to a team motto. Architecture: mock-first, contract-driven FastAPI owns the canonical Pydantic contracts and generates the OpenAPI schema. APP_MODE=mock wires in an in-memory repository and deterministic synthetic fixtures, so the entire loop- three vendor calls, quote generation, negotiation, recommendation- runs without a single external API key or account. That decision mattered enormously for a 24-hour build: nobody was blocked waiting on ElevenLabs quota or an OpenAI key while the demo loop got proven out.&lt;/p&gt;

&lt;p&gt;The frontend's rule, spelled out in the repo's AGENTS.md, was strict: one API client, types generated from FastAPI's OpenAPI schema via openapi-typescript, never handwritten parallel domain models. Any contract change meant: update the Pydantic models and backend tests, re-export the OpenAPI JSON, regenerate the TypeScript types, update call sites, run the full check suite, then get sign-off from both the backend and frontend owners. That discipline is what let four people build in parallel without four different mental models of what a Quote looks like. Four conversation-design requirements, made visible. Because "moving services negotiator" is a fairly obvious vertical for a hackathon like this, differentiation had to come from execution depth, not novelty. One deliberate choice: build a visible checklist in the UI for the four things the challenge brief actually cared about in a voice agent.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn0vryqppbgp47efykgyu.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fn0vryqppbgp47efykgyu.png" alt=" " width="799" height="630"&gt;&lt;/a&gt;&lt;br&gt;
Disclosure: every call opens by stating it's an AI calling on the customer's behalf, and answers "are you a robot?" honestly. Friction survival: at least one call has to survive real friction (hold music, a rushed dispatcher) and land a structured callback instead of a dead end. The honesty line: the agent can use a real verified competing quote as leverage, but must never invent inventory or fabricate a competitor's number. Structured close: no call ends on a vague answer; every call closes as an itemized quote, a callback commitment, or a documented decline.&lt;/p&gt;

&lt;p&gt;We built this as a small four-item card that lights up based on real call outcome data rather than being pure set dressing; two items are static guardrail statements (things baked into every synthetic call by design), and two are computed live from each call's actual outcome type. Honest &amp;gt; flashy, especially when honesty is literally what the requirement is testing for. Hidden fees and red flags, from the schema up. The vendor persona design does a lot of work here. Three synthetic vendors, three distinct negotiation styles: a transparent one that itemizes everything up front, a cheap-headline one that reveals fees only under direct questioning, and a premium one that moves meaningfully when shown a better verified quote.&lt;/p&gt;

&lt;p&gt;On the data side, every fee line item on a quote carries a disclosed_upfront: boolean. Any fee where that's false is, by definition, a hidden fee caught after the fact, which meant the frontend didn't need any fuzzy logic to surface this, just a filter. Red flags work the same way: a red_flags: string[] field lives directly on both the quote and the final recommendation ranking, populated by the backend's own rules (e.g. a quote 30%+ below the median gets flagged, never presented as a clean win without a caveat). The mid-build pivot: Vite/React to Lovable Partway through, after a role swap put me on frontend ownership, we shipped a fully working Vite + React + TypeScript implementation of the loop, intake, confirm, calls, negotiate, report, all wired to the real mock backend, with hidden-fee and red-flag surfacing, a full JobSpec review screen, and test coverage for the new logic. It passed the full CI-equivalent check (Ruff, pytest, OpenAPI export, typecheck, Vitest, production build) and went up as a PR.&lt;/p&gt;

&lt;p&gt;Then the team decided to move the frontend to Lovable. That decision came with a real lesson in the cost of a mid-build tool switch:&lt;/p&gt;

&lt;p&gt;Rebuilding on hearsay contracts. Lovable's agent can't run openapi-typescript against a live backend the way our Vite setup could; it works from whatever you tell it. My first pass described the API contract from memory, and it quietly diverged from the real schema in three ways: money fields (original_total, negotiated_total, deposit) are serialized by FastAPI as decimal strings, not JSON numbers; stairs is a numeric count, not a boolean; and version fields are the literal string "1.0", not a number. All three got caught and fixed, but only because we'd separately verified the real schema earlier in the Vite build; without that, they'd have shipped wrong. Verifying against real payloads, not assumptions. When we needed the shape of a call record for the conversation-design checklist, I initially had Lovable invent field names based on partial information. The actual backend owner sent a real JSON payload back, and the real shape looked meaningfully different: a full Vendor object nested in multiple places instead of a {id, name} shorthand, a status field, and started_at/completed_at timestamps. Type-correcting against a real payload took minutes; building on the wrong assumption and finding out later would have cost a lot more. Unrelated git histories. Lovable auto-provisions its own GitHub repo per project with its own root commit. Pushing that as a new branch onto the real repository worked at the git level, but GitHub's compare view came back with "there isn't anything to compare, main and this branch are entirely different commit histories." The fix was mechanical once diagnosed: branch from the real main, copy the Lovable-exported files into the correct directory, commit fresh on top of the real history, and push that instead. A five-minute problem once you know unrelated-history diffs are a known GitHub limitation, a confusing one if you don't. CORS and reachability. A cloud-hosted frontend can't call &lt;a href="http://127.0.0.1:8000" rel="noopener noreferrer"&gt;http://127.0.0.1:8000&lt;/a&gt;; that address only means something on the machine running it. Anything built in a hosted tool needs either a publicly reachable backend or a tunnel, and the backend's CORS policy needs to explicitly allow the new frontend's origin. Easy to forget when you've been developing against localhost for the whole build.&lt;/p&gt;

&lt;p&gt;None of these were hard problems. All of them were the direct cost of losing the single source of truth (a generated, verified contract) the moment the frontend moved to a tool that couldn't generate types from the live API itself. The fix in every case was the same instinct: stop guessing, go get the real payload, verify against it. What we'd tell another team doing this: Freeze your contracts before writing feature code, and freeze them first between backend and frontend, not last. Build the ugliest possible version of the full loop before touching visual polish. A working call beats a polished interface, every time. If you have to build in a tool that can't generate types from your live API, verify field-by-field against a real payload before you build on top of assumed ones; the difference between "grounded in the generated contract" and "hand-typed from memory" is exactly the gap where subtle, expensive bugs live.&lt;/p&gt;

&lt;p&gt;VeraMove is a hackathon starter, not a production product, mock mode only, synthetic data throughout, no real vendor calls. But the loop it proves one locked spec, three comparable quotes, one negotiation backed by real leverage, one evidence-linked recommendation is the whole idea, and getting that loop working end to end, twice, in two different frontend stacks, taught us more about contract discipline than either build alone would have.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm1zye2ozq0mh8sk6c60z.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fm1zye2ozq0mh8sk6c60z.png" alt=" " width="800" height="579"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Check out the code:&lt;/strong&gt; &lt;a href="https://github.com/zukhriddingit/VeraMove" rel="noopener noreferrer"&gt;github.com/zukhriddingit/VeraMove&lt;/a&gt;. Clone it, run python scripts/&lt;a href="http://bootstrap.py" rel="noopener noreferrer"&gt;bootstrap.py&lt;/a&gt;, then python scripts/&lt;a href="http://dev.py" rel="noopener noreferrer"&gt;dev.py&lt;/a&gt;, and you'll have the full loop running locally in about five minutes no credentials needed&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
