<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Aria Kovac</title>
    <description>The latest articles on DEV Community by Aria Kovac (@ariakovac).</description>
    <link>https://dev.to/ariakovac</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4009441%2Fc47eace3-e41c-4e7d-b333-9829f76dc6c1.png</url>
      <title>DEV Community: Aria Kovac</title>
      <link>https://dev.to/ariakovac</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ariakovac"/>
    <language>en</language>
    <item>
      <title>Beyond Chatbots: A Decision Workflow for Getting Everyone to Kickoff</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Tue, 01 Sep 2026 07:06:44 +0000</pubDate>
      <link>https://dev.to/ariakovac/beyond-chatbots-a-decision-workflow-for-getting-everyone-to-kickoff-1p9l</link>
      <guid>https://dev.to/ariakovac/beyond-chatbots-a-decision-workflow-for-getting-everyone-to-kickoff-1p9l</guid>
      <description>&lt;p&gt;Most match-day plans do not fail because people are lazy.&lt;/p&gt;

&lt;p&gt;They fail because the group chat is trying to use a chatbot answer where it actually needs a decision workflow.&lt;/p&gt;

&lt;p&gt;That difference matters for World Cup-scale football days in New York and New Jersey. The 2026 tournament has already concluded, but its New York/New Jersey match days are still a useful retrospective case study. FIFA scheduled the final for New York/New Jersey on July 19, 2026, and NJ Transit’s 2026 World Cup guidance described eight MetLife Stadium match days, including World Cup-specific rail rules between Penn Station New York and Secaucus Junction in the four hours before kickoff.&lt;/p&gt;

&lt;p&gt;Those details are not decoration. They are the whole problem.&lt;/p&gt;

&lt;p&gt;A fan leaving Queens is not solving the same route as a friend leaving Brooklyn. Someone coming from Manhattan may care about Penn Station timing. Someone already on the New Jersey side may not want to backtrack into the city just because the group chat picked a familiar meeting point. And after the match, the plan changes again, because everybody leaving together at the same time is usually the least realistic version of events.&lt;/p&gt;

&lt;p&gt;A normal chatbot can answer a question like, "How do we get to the stadium?"&lt;/p&gt;

&lt;p&gt;A better match-day assistant should help a group decide.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Wrong Object: One Perfect Recommendation
&lt;/h2&gt;

&lt;p&gt;The weak version of AI planning gives one confident answer:&lt;/p&gt;

&lt;p&gt;"Leave early, meet near the stadium, and take transit."&lt;/p&gt;

&lt;p&gt;That is not wrong. It is just not enough.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNmRmYmJjY2Q2ZDFmYzkzNjQ0OTdjNGVmMjBhYTQ1MWRfanpXTDE4Rk96WDB0WEJOcllzelJlTHQxbmtBYkZ4SmVfVG9rZW46UTdrZGJzRktrb3RuRnl4RGdmYWNhaXhKbnRiXzE3ODgyNDYxMDY6MTc4ODI0OTcwNl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNmRmYmJjY2Q2ZDFmYzkzNjQ0OTdjNGVmMjBhYTQ1MWRfanpXTDE4Rk96WDB0WEJOcllzelJlTHQxbmtBYkZ4SmVfVG9rZW46UTdrZGJzRktrb3RuRnl4RGdmYWNhaXhKbnRiXzE3ODgyNDYxMDY6MTc4ODI0OTcwNl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="1170" height="780"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It ignores where people start. It ignores who can leave work early. It ignores who hates transfers, who needs food first, who has the tickets, and who will be impossible to reach once the cell network gets noisy.&lt;/p&gt;

&lt;p&gt;I see the same shape in support work all the time. The failure is not always bad information. Sometimes the failure is that the information never became an assignable action.&lt;/p&gt;

&lt;p&gt;For match day, the useful object is not a recommendation. It is a small decision packet:&lt;/p&gt;

&lt;p&gt;Context. Options. Reason. Action.&lt;/p&gt;

&lt;p&gt;That is it. Not model architecture. Not a giant itinerary. Just enough structure to move a group from "what should we do?" to "we are doing this."&lt;/p&gt;

&lt;h2&gt;
  
  
  The iMessage Version
&lt;/h2&gt;

&lt;p&gt;Here is the kind of flow I would want in a group chat before kickoff:&lt;/p&gt;

&lt;p&gt;Context: Four people. One in Astoria, one in Bushwick, one near Harlem, one in Jersey City. In this example, kickoff is at 6:00 PM. Two people can leave early. Nobody wants to drive. The group wants food before entering the stadium, but not a long sit-down dinner.&lt;/p&gt;

&lt;p&gt;Option 1: Everyone meets near Penn Station New York, then travels together toward Secaucus and the stadium.&lt;/p&gt;

&lt;p&gt;Option 2: Split by side of the river. The NYC group meets near Penn Station. The Jersey City person meets the group closer to Secaucus. Food happens before the final stadium leg, not near each person's apartment.&lt;/p&gt;

&lt;p&gt;Option 3: Skip the pre-stadium meet-up. Everyone enters separately, then meets at a fixed section or landmark inside after security.&lt;/p&gt;

&lt;p&gt;Reason: Option 2 gives the group one shared checkpoint without forcing the New Jersey person to travel backward into Manhattan. It also avoids making the latest person control everyone else's arrival.&lt;/p&gt;

&lt;p&gt;Action: Pick Option 2. NYC group leaves first. Jersey City person confirms arrival separately. If anyone misses the checkpoint, they switch to Option 3 instead of restarting the whole plan.&lt;/p&gt;

&lt;p&gt;That is the level of structure I want from city AI. For this kind of local decision, &lt;a href="https://app.karpo.ai/how-it-works" rel="noopener noreferrer"&gt;Karpo for everyday city decisions&lt;/a&gt; makes more sense than a generic chatbot because the useful input is not just a question. It starts with mood, neighborhood, budget, and who is coming with you, then turns that context into places, events, and plans.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMGQxZDVmNWUxOGZjYjM3Y2EwYjI3ZGIwNjI4M2FjM2FfQ3ZTMU5lN0lZeU8wc0JtNDI2eUFBY0VvM01YWTNxeXVfVG9rZW46R3cwdWI1alBDbzlJT3J4SHVYSWNBdWhGbktYXzE3ODgyNDYxMDY6MTc4ODI0OTcwNl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMGQxZDVmNWUxOGZjYjM3Y2EwYjI3ZGIwNjI4M2FjM2FfQ3ZTMU5lN0lZeU8wc0JtNDI2eUFBY0VvM01YWTNxeXVfVG9rZW46R3cwdWI1alBDbzlJT3J4SHVYSWNBdWhGbktYXzE3ODgyNDYxMDY6MTc4ODI0OTcwNl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="1382" height="639"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Not an official event guide. Not a tournament partner. Just a better way to turn local constraints into a plan people can actually follow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Before Kickoff: Solve Convergence
&lt;/h2&gt;

&lt;p&gt;Before kickoff, the main problem is not the stadium. It is convergence.&lt;/p&gt;

&lt;p&gt;Where does the group become a group?&lt;/p&gt;

&lt;p&gt;For NYC football fans, that question has layers. A plan that sounds simple from Manhattan may be annoying from Queens. A plan that looks fast on a map may fail once you add event-day transit rules, transfer anxiety, bag checks, weather, and the one friend who always sends "almost there" from twenty minutes away.&lt;/p&gt;

&lt;p&gt;I would keep pre-match decisions narrow:&lt;/p&gt;

&lt;p&gt;Pick one shared checkpoint.&lt;/p&gt;

&lt;p&gt;Pick one fallback checkpoint.&lt;/p&gt;

&lt;p&gt;Pick a latest arrival time.&lt;/p&gt;

&lt;p&gt;Decide who carries the tickets or confirms the ticket transfer.&lt;/p&gt;

&lt;p&gt;Decide whether food is before transit, near the checkpoint, or after entry.&lt;/p&gt;

&lt;p&gt;Everything else is noise.&lt;/p&gt;

&lt;p&gt;If an AI tool gives seven restaurant choices, three transit routes, and a motivational paragraph, it has not helped yet. It has created more surface area for disagreement.&lt;/p&gt;

&lt;p&gt;The best output is usually two to four options with trade-offs.&lt;/p&gt;

&lt;h2&gt;
  
  
  During the Match: Plan for Drift
&lt;/h2&gt;

&lt;p&gt;During the match, the plan should become smaller, not larger.&lt;/p&gt;

&lt;p&gt;People lose signal. Someone goes for water. Someone is late. Someone's phone battery drops to nine percent. Someone wants to stay in the seat while everyone else wants to move.&lt;/p&gt;

&lt;p&gt;A good decision workflow handles that without drama:&lt;/p&gt;

&lt;p&gt;If separated before kickoff, meet at the section.&lt;/p&gt;

&lt;p&gt;If separated during the match, do not move unless there is a clear reason.&lt;/p&gt;

&lt;p&gt;If signal fails, use the last agreed checkpoint.&lt;/p&gt;

&lt;p&gt;If the group needs to split, name the next reconnection time.&lt;/p&gt;

&lt;p&gt;This is not glamorous AI. It is just coordination under stress. But that is exactly where generic chatbots feel thin. They can suggest. They cannot always commit the group to a next step.&lt;/p&gt;

&lt;h2&gt;
  
  
  After the Match: Do Not Pretend Everyone Leaves Together
&lt;/h2&gt;

&lt;p&gt;The after-match plan deserves its own decision.&lt;/p&gt;

&lt;p&gt;People often plan the arrival carefully and then treat the exit as an afterthought. That is backwards. The exit has more fatigue, less patience, weaker phone batteries, and heavier crowd pressure.&lt;/p&gt;

&lt;p&gt;NJ Transit’s 2026 World Cup match-day guidance treated post-match travel as its own operating window, with targeted service adjustments after matches. That is the kind of planning fact a match-day assistant should respect. Even if you are not using those exact past dates, the principle holds for any major match day: post-event travel is a different decision from pre-event travel.&lt;/p&gt;

&lt;p&gt;I would give the group three exit options before the match starts:&lt;/p&gt;

&lt;p&gt;Option A: Leave together immediately.&lt;/p&gt;

&lt;p&gt;Option B: Wait out the first rush at a nearby safe, agreed location.&lt;/p&gt;

&lt;p&gt;Option C: Split by destination and stop pretending one route serves everyone.&lt;/p&gt;

&lt;p&gt;Then pick the default before kickoff.&lt;/p&gt;

&lt;p&gt;The point is not to predict the entire evening. The point is to remove one tired argument from the end of it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Developers Should Notice
&lt;/h2&gt;

&lt;p&gt;The useful product lesson here is not "add AI to event planning."&lt;/p&gt;

&lt;p&gt;It is that some consumer AI tasks should not be shaped like Q&amp;amp;A.&lt;/p&gt;

&lt;p&gt;A match-day group does not need an answer. It needs a decision object with ownership, timing, fallbacks, and local constraints.&lt;/p&gt;

&lt;p&gt;For a city decision tool, I would rather see:&lt;/p&gt;

&lt;p&gt;Two to four realistic options.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNmJlMjFhMTg4Yjc0ZTg1MmE5MmU0NjFlYjM5YzM2ZmVfSVp6emVyWktxckMxQ1JxbWszNlNnZ0lRbkFwTWd1MnFfVG9rZW46T09wTmIzZ3Nmb2tRQ0t4c0dMTmNjaEY0bnpnXzE3ODgyNDYxMDY6MTc4ODI0OTcwNl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNmJlMjFhMTg4Yjc0ZTg1MmE5MmU0NjFlYjM5YzM2ZmVfSVp6emVyWktxckMxQ1JxbWszNlNnZ0lRbkFwTWd1MnFfVG9rZW46T09wTmIzZ3Nmb2tRQ0t4c0dMTmNjaEY0bnpnXzE3ODgyNDYxMDY6MTc4ODI0OTcwNl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="800" height="534"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A short reason for each.&lt;/p&gt;

&lt;p&gt;A clear default.&lt;/p&gt;

&lt;p&gt;A fallback if the group misses the plan.&lt;/p&gt;

&lt;p&gt;A source-of-truth warning when event details, transit schedules, or venue policies may have changed.&lt;/p&gt;

&lt;p&gt;That last part matters. Any AI product touching live city plans should know when to stop being confident. Match times, service advisories, bag policies, street closures, and weather can change. A good assistant should push users toward current official sources when the details are operational.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Shift
&lt;/h2&gt;

&lt;p&gt;The future of everyday AI is not just better chat.&lt;/p&gt;

&lt;p&gt;It is better small decisions.&lt;/p&gt;

&lt;p&gt;Where should we meet? Who goes first? Which option works for the person coming from the farthest borough? What happens if the train is delayed? What do we do after the match, when everyone is tired and nobody wants to reopen the discussion?&lt;/p&gt;

&lt;p&gt;That is where city AI can become useful without pretending to be magic.&lt;/p&gt;

&lt;p&gt;World Cup-scale match days make the problem obvious, but the pattern is ordinary. Concerts, playoff games, airport pickups, late dinners, weekend plans, visiting friends, bad weather, different budgets, different neighborhoods.&lt;/p&gt;

&lt;p&gt;A chatbot gives an answer.&lt;/p&gt;

&lt;p&gt;A decision workflow gets the group moving.&lt;/p&gt;

&lt;p&gt;On a crowded match day, that difference is not theoretical. It is whether everyone makes it to kickoff still speaking to each other.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Final Gate&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Risk lowered&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;“It starts with mood, neighborhood, budget, and who is coming with you...”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Event timing corrected&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;“The 2026 tournament has already concluded...”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NJ Transit boundary narrowed&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;“World Cup-specific rail rules between Penn Station New York and Secaucus Junction...”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No official affiliation claim&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;“Not an official event guide. Not a tournament partner.”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assigned anchor used once&lt;/td&gt;
&lt;td&gt;pass&lt;/td&gt;
&lt;td&gt;“Karpo for everyday city decisions”&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>productivity</category>
      <category>ux</category>
    </item>
    <item>
      <title>Natural’s $30M Raise Shows Why Agentic Payments Are the Missing Runtime for AI Agents</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Tue, 21 Jul 2026 09:02:11 +0000</pubDate>
      <link>https://dev.to/ariakovac/naturals-30m-raise-shows-why-agentic-payments-are-the-missing-runtime-for-ai-agents-4fa0</link>
      <guid>https://dev.to/ariakovac/naturals-30m-raise-shows-why-agentic-payments-are-the-missing-runtime-for-ai-agents-4fa0</guid>
      <description>&lt;p&gt;I do not worry much when an AI agent writes a bad draft. I worry when it can hire a vendor, agree to a price, and then has no safe way to pay.&lt;/p&gt;

&lt;p&gt;That is the real signal behind Natural’s new funding round. On July 20, 2026, &lt;a href="https://www.natural.com/blog/natural-series-a" rel="noopener noreferrer"&gt;Natural announced&lt;/a&gt; a \$30 million Series A led by Kirsten Green at Forerunner, bringing its total funding to over \$40 million. The company says it is building the foundational payments stack for AI agents, with live products including wallets, vaults, pay, request, transfer, and connect.&lt;/p&gt;

&lt;p&gt;The obvious headline is “Natural wants to take on Stripe.”&lt;/p&gt;

&lt;p&gt;The more useful developer framing is this: once agents move from recommending actions to executing them, payments become a runtime problem.&lt;/p&gt;

&lt;p&gt;Today’s payment systems assume a human is somewhere nearby. A person enters card details. A person approves a bank transfer. A person knows why the invoice exists. A person can explain the dispute later.&lt;/p&gt;

&lt;p&gt;Agent workflows break that assumption.&lt;/p&gt;

&lt;p&gt;Imagine a logistics agent that finds a freight vendor, compares quotes, negotiates timing, and prepares a shipment. The moment money needs to move, the agent has to stop and ask a human to finish the transaction. That is not a small UX issue. It is the boundary between automation and actual economic agency.&lt;/p&gt;

&lt;p&gt;A payment tool for agents cannot just be a checkout button with a prettier API. It needs policy, identity, observability, and failure handling.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZjZlYzQ4ODRhODlkMThiNDg3NmU1ZDY1MGE0YTgxOTRfSmlUS3pLRnN5R0FOMWJvNFlOM05zaG85aFNQVzNYTHNfVG9rZW46QUtjUWJ3dFBtb0hGMmd4TnVMR2NkeDVubnZlXzE3ODQ2MjQwNTQ6MTc4NDYyNzY1NF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZjZlYzQ4ODRhODlkMThiNDg3NmU1ZDY1MGE0YTgxOTRfSmlUS3pLRnN5R0FOMWJvNFlOM05zaG85aFNQVzNYTHNfVG9rZW46QUtjUWJ3dFBtb0hGMmd4TnVMR2NkeDVubnZlXzE3ODQ2MjQwNTQ6MTc4NDYyNzY1NF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Screenshot from Natural" width="800" height="589"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Screenshot from &lt;a href="https://www.natural.com/" rel="noopener noreferrer"&gt;Natural&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If I were designing an agentic payment flow, I would want the payment layer to answer seven questions before funds move:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Who authorized this agent to spend?&lt;/li&gt;
&lt;li&gt;What budget, rail, merchant, and category limits apply?&lt;/li&gt;
&lt;li&gt;Is the counterparty an agent, a business, or a human?&lt;/li&gt;
&lt;li&gt;What evidence connects the payment to the original task?&lt;/li&gt;
&lt;li&gt;Does this payment require human approval?&lt;/li&gt;
&lt;li&gt;What happens if the agent made the wrong decision?&lt;/li&gt;
&lt;li&gt;Can support, finance, or compliance reconstruct the transaction later?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That last point is where my support-engineering brain gets loud.&lt;/p&gt;

&lt;p&gt;A bad agentic payment system does not only create fraud risk. It creates explanation debt. Someone will eventually ask why an agent paid a vendor, why it chose that amount, why it skipped another option, or why a refund request should be accepted. If the payment layer cannot show the trace, the problem becomes a human investigation.&lt;/p&gt;

&lt;p&gt;Natural’s own product language points in this direction. Its &lt;a href="https://www.natural.com/pay" rel="noopener noreferrer"&gt;Pay page&lt;/a&gt; says Natural handles orchestration, ledgering, routing, compliance, risk, and disputes. Its earlier &lt;a href="https://www.natural.com/blog/agentic-payments-memo" rel="noopener noreferrer"&gt;agentic payments memo&lt;/a&gt; talks about controllable wallets, transaction-level rules, authorization tools, and observability.&lt;/p&gt;

&lt;p&gt;Those are exactly the boring primitives that make autonomous systems survivable.&lt;/p&gt;

&lt;p&gt;Stripe is not ignoring the same problem. &lt;a href="https://docs.stripe.com/agentic-commerce" rel="noopener noreferrer"&gt;Stripe’s agentic commerce docs&lt;/a&gt; describe flows where agents help buyers browse products, manage carts, and complete purchases. The docs also reference protocols such as UCP, ACP, MPP, and x402 across different seller and agent flows. Stripe has also written about &lt;a href="https://stripe.com/blog/supporting-additional-payment-methods-for-agentic-commerce" rel="noopener noreferrer"&gt;Shared Payment Tokens&lt;/a&gt;, where agents can initiate payments with scoped credentials instead of exposing the underlying payment method.&lt;/p&gt;

&lt;p&gt;That tells me this is not one startup’s narrative. It is an infrastructure race.&lt;/p&gt;

&lt;p&gt;There are also stablecoin-focused players such as &lt;a href="https://skyfire.xyz/overview-ai-agents-and-payments/" rel="noopener noreferrer"&gt;Skyfire&lt;/a&gt;, which describes itself as identity and payments infrastructure for AI agents. The stablecoin angle makes sense for some machine-speed payments, especially API access, global settlement, and microtransactions. But I would be careful about assuming stablecoins solve the whole problem. Enterprises still care about bank rails, reconciliation, compliance workflows, and accounting systems that do not disappear because an agent can hold a wallet.&lt;/p&gt;

&lt;p&gt;So the question is not simply Natural vs Stripe.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYTUwYTU3YzcwYzViYjQ0MzgzYWUyMzE5YmRhNDQ1MTFfdzRGUUpZQ05EVHlrZnd6bjZHdWc2c1ZXRUJTajRreHFfVG9rZW46R1RCRGJDdmRYb20yZmd4OUZQSGNkcVd6bkJmXzE3ODQ2MjQwNTQ6MTc4NDYyNzY1NF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYTUwYTU3YzcwYzViYjQ0MzgzYWUyMzE5YmRhNDQ1MTFfdzRGUUpZQ05EVHlrZnd6bjZHdWc2c1ZXRUJTajRreHFfVG9rZW46R1RCRGJDdmRYb20yZmd4OUZQSGNkcVd6bkJmXzE3ODQ2MjQwNTQ6MTc4NDYyNzY1NF9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Screenshot from Stripe" width="800" height="505"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Screenshot from &lt;a href="https://stripe.com/" rel="noopener noreferrer"&gt;Stripe&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The question is what kind of payment primitive agent developers will actually use.&lt;/p&gt;

&lt;p&gt;For small agent workflows, maybe a scoped card token is enough. For marketplace agents, maybe the hard problem is merchant onboarding. For business operations agents, ACH and bank-account workflows may matter more. For agent-to-agent API markets, wallets and real-time settlement may be cleaner.&lt;/p&gt;

&lt;p&gt;The dangerous mistake would be treating payments as the final step after the agent is already built.&lt;/p&gt;

&lt;p&gt;Payments should be designed as part of the agent architecture:&lt;/p&gt;

&lt;p&gt;The agent should never see raw financial credentials.&lt;/p&gt;

&lt;p&gt;Every payment action should carry a task ID, policy decision, counterparty identity, and audit trace.&lt;/p&gt;

&lt;p&gt;Approval rules should be deterministic, not hidden inside a prompt.&lt;/p&gt;

&lt;p&gt;Failed payments should be typed errors the agent can handle.&lt;/p&gt;

&lt;p&gt;Disputes should attach back to the agent’s reasoning and tool history.&lt;/p&gt;

&lt;p&gt;Spend limits should exist at the organization, wallet, agent, and transaction level.&lt;/p&gt;

&lt;p&gt;The payment layer should make “no” cheap.&lt;/p&gt;

&lt;p&gt;That last line matters. Most demos optimize for agents saying yes quickly. Production systems need agents that can stop cleanly, escalate politely, and explain why money did not move.&lt;/p&gt;

&lt;p&gt;Natural’s \$30 million raise is interesting because it puts a spotlight on the least glamorous part of agent autonomy.&lt;/p&gt;

&lt;p&gt;Not the model.&lt;/p&gt;

&lt;p&gt;Not the browser.&lt;/p&gt;

&lt;p&gt;Not the conversation UI.&lt;/p&gt;

&lt;p&gt;The money movement.&lt;/p&gt;

&lt;p&gt;And money movement changes the trust model.&lt;/p&gt;

&lt;p&gt;When an agent summarizes an article badly, you fix the article.&lt;/p&gt;

&lt;p&gt;When an agent pays the wrong vendor, you need authorization logs, dispute rules, transaction history, counterparty identity, refund paths, and probably a tired person in support trying to make sense of the whole thing.&lt;/p&gt;

&lt;p&gt;That is why agentic payments feel like a real developer topic, not just fintech gossip. If agents are going to act on behalf of people and businesses, they need financial infrastructure built for software actors.&lt;/p&gt;

&lt;p&gt;The best agentic payment layer will not be the one that makes agents spend fastest.&lt;/p&gt;

&lt;p&gt;It will be the one that makes every payment explainable after the fact.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>programming</category>
      <category>automation</category>
    </item>
    <item>
      <title>Self-Evolving Agents Are Not AutoGPT With Better Memory</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Mon, 20 Jul 2026 10:28:49 +0000</pubDate>
      <link>https://dev.to/ariakovac/self-evolving-agents-are-not-autogpt-with-better-memory-34a4</link>
      <guid>https://dev.to/ariakovac/self-evolving-agents-are-not-autogpt-with-better-memory-34a4</guid>
      <description>&lt;p&gt;I do not trust an AI agent because it can run for a long time.&lt;/p&gt;

&lt;p&gt;I start trusting it when I can see what changed after it failed.&lt;/p&gt;

&lt;p&gt;That is the useful way to read the current wave of self-evolving agents. The phrase sounds dramatic, but the engineering question is very plain: can an agent turn experience into a durable, verified improvement without quietly making itself worse?&lt;/p&gt;

&lt;p&gt;As of July 2026, several threads are converging. PaperAgent’s BestHub digest frames the field around model-centric evolution, environment-centric evolution, and model-environment co-evolution. The arXiv survey &lt;a href="https://arxiv.org/abs/2507.21046" rel="noopener noreferrer"&gt;A Survey of Self-Evolving Agents&lt;/a&gt; organizes the problem around what evolves, when it evolves, and how that evolution is guided. Another survey, &lt;a href="https://arxiv.org/abs/2508.07407" rel="noopener noreferrer"&gt;A Comprehensive Survey of Self-Evolving AI Agents&lt;/a&gt;, defines self-evolving agents as systems that systematically optimize internal components through environment interaction while preserving safety and performance.&lt;/p&gt;

&lt;p&gt;That last clause is the part I care about.&lt;/p&gt;

&lt;p&gt;“Self-evolving” does not mean “the agent keeps trying.” It means the agent has a loop where feedback changes something reusable: a prompt, memory, tool, workflow, evaluator, policy, or model component.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Smallest Useful Architecture
&lt;/h2&gt;

&lt;p&gt;A practical self-evolving agent architecture needs five pieces:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Task / User Goal
      |
      v
Agent Runner  ---&amp;gt; Tools / Environment
      |                 |
      v                 v
Trace Store &amp;lt;--- Results / Failures
      |
      v
Evaluator / Verifier
      |
      v
Evolution Planner
      |
      v
Proposal Queue ---&amp;gt; Approval Gate ---&amp;gt; Versioned Agent State
                         |
                         v
                      Rollback
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important part is not the agent runner. We already have plenty of those.&lt;/p&gt;

&lt;p&gt;The important part is the path from failure to a versioned change.&lt;/p&gt;

&lt;p&gt;If the agent fails at a task, the system should capture the trace, score it, explain the weakness, propose an update, test that update, and only then retain it. Without that retention step, you have retry logic. Without the verifier, you have vibes. Without rollback, you have an incident waiting for a polite calendar invite.&lt;/p&gt;

&lt;p&gt;Here is the pseudocode version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;state = load_agent_state()

while True:
    task = receive_task()
    trace = run_agent(task, state)

    score = evaluate(trace, task.success_criteria)

    if score.passed:
        maybe_store_success_pattern(trace, state)
        continue

    weakness = diagnose_failure(trace)
    proposal = propose_state_update(
        weakness=weakness,
        mutable_targets=["prompt", "memory", "tool_policy", "workflow_graph"]
    )

    test_result = verify_update(proposal, regression_suite=state.tests)

    if test_result.passed and approval_gate(proposal):
        state = commit_versioned_update(state, proposal)
    else:
        keep_state_unchanged()
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is the boring shape of a real self-evolving agent architecture. Boring is good here. Boring means someone can debug it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYTVmMjMxM2Q2MjU0MGU4NzRjNzNjNTQ3YzM5MWRkZDlfTTg4aThRc1dMMjNUdWQyQUdQM1RFMjdQbFptQ2VuSmlfVG9rZW46T2VoNWI3eUtmb2diMkF4MGxSc2NNWkt4bjNjXzE3ODQ1NDMwNDU6MTc4NDU0NjY0NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYTVmMjMxM2Q2MjU0MGU4NzRjNzNjNTQ3YzM5MWRkZDlfTTg4aThRc1dMMjNUdWQyQUdQM1RFMjdQbFptQ2VuSmlfVG9rZW46T2VoNWI3eUtmb2diMkF4MGxSc2NNWkt4bjNjXzE3ODQ1NDMwNDU6MTc4NDU0NjY0NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Photo by Rahul Mishra on Unsplash" width="870" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo by &lt;a href="https://unsplash.com/@rahuulmiishra" rel="noopener noreferrer"&gt;Rahul Mishra&lt;/a&gt; on &lt;a href="https://unsplash.com/" rel="noopener noreferrer"&gt;Unsplash&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  AutoGPT vs Self-Evolving Agents
&lt;/h2&gt;

&lt;p&gt;AutoGPT was important because it made autonomous agents feel tangible. The current &lt;a href="https://github.com/significant-gravitas/autogpt" rel="noopener noreferrer"&gt;AutoGPT repository&lt;/a&gt; describes the project as a platform to create, deploy, and manage continuous AI agents that automate workflows.&lt;/p&gt;

&lt;p&gt;But AutoGPT-style autonomy and self-evolution are not the same thing.&lt;/p&gt;

&lt;p&gt;An autonomous agent can plan tasks, call tools, and chain actions. A self-evolving agent changes the system that will handle the next task.&lt;/p&gt;

&lt;p&gt;The difference is persistence plus validation.&lt;/p&gt;

&lt;p&gt;If an agent searches the web, writes a plan, fails, and tries again, that is autonomy.&lt;/p&gt;

&lt;p&gt;If it notices that its tool-selection policy caused the failure, proposes a routing change, tests that change against previous tasks, stores the update, and can roll it back later, that is self-evolution.&lt;/p&gt;

&lt;p&gt;This is why I would not describe self-evolving agents as “AutoGPT with better memory.” Memory is only one mutable component. The deeper shift is that the agent scaffold itself becomes optimizable.&lt;/p&gt;

&lt;p&gt;A July 2026 survey, &lt;a href="https://arxiv.org/abs/2607.13104" rel="noopener noreferrer"&gt;Self-Improvements in Modern Agentic Systems&lt;/a&gt;, frames a modern agent as a foundation model coupled with prompts, memory, tools, and control logic. It also treats self-improvement as an update operator that can commit changes to model parameters or scaffold components. That is a much cleaner mental model.&lt;/p&gt;

&lt;p&gt;The model is not the whole agent. But the scaffold matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should Actually Evolve?
&lt;/h2&gt;

&lt;p&gt;For most developers, model-weight evolution is not the first place to start. It is expensive, risky, and hard to evaluate.&lt;/p&gt;

&lt;p&gt;The practical starting points are scaffold components:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Prompts: rewrite instructions based on repeated failure modes.&lt;/li&gt;
&lt;li&gt;Memory: decide what to retain, merge, forget, or distrust.&lt;/li&gt;
&lt;li&gt;Tool policy: change when and how tools are selected.&lt;/li&gt;
&lt;li&gt;Workflow graph: alter the order of planner, executor, critic, verifier, and human review.&lt;/li&gt;
&lt;li&gt;Test set: add new regression cases from real failures.&lt;/li&gt;
&lt;li&gt;Evaluator rubric: improve the judge, but with extra caution.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The evaluator is the most dangerous piece to evolve. If the agent learns to please a weak judge, it may improve the score while degrading the work.&lt;/p&gt;

&lt;p&gt;This is not theoretical. Google Research’s work on &lt;a href="https://research.google/blog/towards-a-science-of-scaling-agent-systems-when-and-why-agent-systems-work/" rel="noopener noreferrer"&gt;scaling agent systems&lt;/a&gt; found that multi-agent systems can help on parallelizable tasks but hurt sequential ones. Their study also reported that independent multi-agent systems can amplify errors badly, while centralized orchestration acts more like a validation bottleneck.&lt;/p&gt;

&lt;p&gt;That is a useful warning for self-evolving systems: more agents do not automatically mean more intelligence. Sometimes they just create more places for a bad assumption to reproduce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Gemini 3 and Agnes Fit
&lt;/h2&gt;

&lt;p&gt;Skywork’s pieces on &lt;a href="https://skywork.ai/blog/ai-agent/gemini-3-ai-agents-2025/" rel="noopener noreferrer"&gt;Gemini 3 for AI agents&lt;/a&gt; and &lt;a href="https://skywork.ai/blog/agnes-ai-why-its-working/" rel="noopener noreferrer"&gt;Agnes AI&lt;/a&gt; are not primary research sources, but they are useful market signals.&lt;/p&gt;

&lt;p&gt;The Gemini 3 article argues for a minimal, auditable agent loop: intent grounding, tool calls, verification, human approval, logging, and metrics. That lines up with the engineering direction Google described when it launched &lt;a href="https://blog.google/products-and-platforms/products/gemini/gemini-3-collection/" rel="noopener noreferrer"&gt;Gemini 3&lt;/a&gt; with stronger reasoning, multimodality, coding, and agentic capabilities.&lt;/p&gt;

&lt;p&gt;The Agnes article is more product-facing, but its “workflows, not widgets” framing is also relevant. A self-evolving agent does not live as a floating chatbot. It needs persistent work surfaces, shared memory, review points, and outputs that other people can inspect.&lt;/p&gt;

&lt;p&gt;In support work, that distinction matters. A clever agent that cannot leave a readable trace is not helpful. It is just another system someone has to reverse-engineer during a bad week.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYzBmOTZmMmRiNThkMTkxN2NkNWRjMGVlZTE3ZTY4NTdfbUNCbkR0QUw0c0ZIQXRBM0gwb1lwVjI3RUo5N2l3WWpfVG9rZW46VjhBOWJidDdjbzBpRER4SUZFdGNPa2Z0blhmXzE3ODQ1NDMwNDU6MTc4NDU0NjY0NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYzBmOTZmMmRiNThkMTkxN2NkNWRjMGVlZTE3ZTY4NTdfbUNCbkR0QUw0c0ZIQXRBM0gwb1lwVjI3RUo5N2l3WWpfVG9rZW46VjhBOWJidDdjbzBpRER4SUZFdGNPa2Z0blhmXzE3ODQ1NDMwNDU6MTc4NDU0NjY0NV9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Screenshot from Gemini" width="1180" height="807"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Screenshot from &lt;a href="https://aistudio.google.com/models/gemini-3" rel="noopener noreferrer"&gt;Gemini&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Five Questions I Would Use Before Calling Anything Self-Evolving
&lt;/h2&gt;

&lt;p&gt;If a product or paper claims self-evolution, I would ask:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What is the mutable object?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Is it the prompt, memory, toolset, workflow, evaluator, policy, model weights, or multi-agent topology?&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What feedback signal drives the change?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Formal tests, environment rewards, human review, LLM judges, user behavior, or the agent’s own self-assessment?&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Who verifies the update?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A deterministic test is stronger than a human vibe check. A human review is stronger than a loose LLM judge. The weakest verifier is the same agent praising its own improvement.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Is the change retained?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If nothing durable changes, it is not evolution. It is a retry loop.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Can it roll back?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A system that can improve itself but cannot undo a bad update is not mature. It is brave in the worst possible way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Technical Frontier
&lt;/h2&gt;

&lt;p&gt;The interesting part of self-evolving agents is not that they might become magical independent researchers next month.&lt;/p&gt;

&lt;p&gt;The interesting part is much more grounded: agent systems are becoming measurable, versioned, and improvable.&lt;/p&gt;

&lt;p&gt;That gives developers a new design target. Instead of asking, “Can I make this agent smarter?”, we can ask:&lt;/p&gt;

&lt;p&gt;Can I make its failure memory better?&lt;/p&gt;

&lt;p&gt;Can I make its tool policy safer?&lt;/p&gt;

&lt;p&gt;Can I make its evaluator harder to fool?&lt;/p&gt;

&lt;p&gt;Can I make its workflow adapt without hiding the change?&lt;/p&gt;

&lt;p&gt;Can I preserve performance on old tasks while improving on new ones?&lt;/p&gt;

&lt;p&gt;That is where self-evolving agent architecture becomes useful as an engineering topic rather than a buzzword.&lt;/p&gt;

&lt;p&gt;Self-evolving agents are not about removing humans from the loop as fast as possible. They are about deciding which parts of the loop can safely learn, under what evidence, with what rollback path.&lt;/p&gt;

&lt;p&gt;The strongest systems will not be the ones that say, “I improved myself.”&lt;/p&gt;

&lt;p&gt;They will be the ones that can show the diff.&lt;/p&gt;

&lt;p&gt;Final Gate: pass on language cleanup. Evidence: the revised article uses English-only terminology such as “self-evolving agents,” “self-evolving agent architecture,” and “autonomous agents.”&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>automation</category>
      <category>coding</category>
    </item>
    <item>
      <title>Claude Code vs Cursor 2.0 vs Codex: AI Coding Tools Are Becoming Control Rooms</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Fri, 17 Jul 2026 09:20:58 +0000</pubDate>
      <link>https://dev.to/ariakovac/claude-code-vs-cursor-20-vs-codex-ai-coding-tools-are-becoming-control-rooms-57fl</link>
      <guid>https://dev.to/ariakovac/claude-code-vs-cursor-20-vs-codex-ai-coding-tools-are-becoming-control-rooms-57fl</guid>
      <description>&lt;p&gt;A year ago, the most common question around AI coding tools was simple: can it write the patch?&lt;/p&gt;

&lt;p&gt;In 2026, the better question is: can it survive the whole loop?&lt;/p&gt;

&lt;p&gt;By “whole loop,” I mean everything that happens after the impressive demo. Choosing the task. Keeping multiple agents visible. Approving or rejecting changes. Managing quota. Testing before CI. Explaining the result to a teammate. Making sure the support queue does not inherit a beautiful, half-tested refactor.&lt;/p&gt;

&lt;p&gt;That is why Codex Micro, Claude Code, Cursor 2.0, Alibaba Cloud’s Qoder CN / Model Studio path, and CircleCI Chunk Sidecars feel connected. They are not the same product. But they point in the same direction: AI coding is moving from a chat box into a full development control layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex Micro Is Not Really About the Keyboard
&lt;/h2&gt;

&lt;p&gt;OpenAI’s Codex Micro is listed at &lt;strong&gt;US\$230&lt;/strong&gt; on the official &lt;a href="https://openai.com/supply/co-lab/work-louder/" rel="noopener noreferrer"&gt;OpenAI Supply Co. page&lt;/a&gt;, so the circulating “\$99 Codex keyboard” framing is not accurate.&lt;/p&gt;

&lt;p&gt;The more interesting part is not the price. It is the product assumption.&lt;/p&gt;

&lt;p&gt;Codex Micro is a small Work Louder collaboration with 13 mechanical switches, a touch sensor, a rotary encoder, a planar joystick, RGB status lights, and physical controls mapped to Codex workflows. OpenAI describes it as a “command center for agentic work,” with controls for things like PR review, debugging, refactoring, accepting or rejecting changes, and adjusting reasoning level.&lt;/p&gt;

&lt;p&gt;That sounds niche because it is niche.&lt;/p&gt;

&lt;p&gt;But the interface idea is not silly. Once developers run more than one coding agent at a time, a plain prompt box starts to feel thin. You need status. You need approval controls. You need to know which agent is waiting, which one is running, and which one has finished. You need a way to steer work without constantly context-switching between chats.&lt;/p&gt;

&lt;p&gt;A serious Codex keyboard review, then, should not ask only whether a US\$230 macro pad is “worth it.” It should ask whether agentic coding is becoming operational enough to need hardware controls at all.&lt;/p&gt;

&lt;p&gt;The answer seems to be yes.&lt;/p&gt;

&lt;p&gt;OpenAI’s &lt;a href="https://help.openai.com/zh-hans-cn/articles/20001106-codex-rate-card" rel="noopener noreferrer"&gt;Codex rate card&lt;/a&gt; points in the same direction. Codex pricing moved to token-based credit accounting in April 2026, with usage tied to input, cached input, and output tokens. The page also points users toward usage panels for remaining credit, purchases, and auto-reload management.&lt;/p&gt;

&lt;p&gt;That is less glamorous than a keyboard. It is also more important.&lt;/p&gt;

&lt;p&gt;If agents run longer, in parallel, and across real repositories, budget control becomes part of the developer workflow. The community joke about a “cyber godfather” watching your quota is funny, but the real signal is governance: agentic coding needs cost visibility, not just clever completions.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DOGViOWMwMzllMTI3YmI1YTM0OTVjZTUyZTNkZDMwMTlfU2NTSEN4S2RtZGpYd3pJTVdiZDJJQXU0ZEVSQ0lDd3lfVG9rZW46RnhaaWJldUExb21hVXV4Uk8zWGM3QnNCbkJiXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DOGViOWMwMzllMTI3YmI1YTM0OTVjZTUyZTNkZDMwMTlfU2NTSEN4S2RtZGpYd3pJTVdiZDJJQXU0ZEVSQ0lDd3lfVG9rZW46RnhaaWJldUExb21hVXV4Uk8zWGM3QnNCbkJiXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Work Louder Codex Micro Keyboard | Custom Mechanical Keycaps" width="1452" height="924"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo from &lt;a href="https://openai.com/supply/co-lab/work-louder/" rel="noopener noreferrer"&gt;OpenAI&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Code Is Becoming a Work Surface
&lt;/h2&gt;

&lt;p&gt;Claude Code is already more than a terminal assistant. Anthropic’s &lt;a href="https://code.claude.com/docs/en/overview" rel="noopener noreferrer"&gt;Claude Code overview&lt;/a&gt; describes it as an agentic coding tool that reads your codebase, edits files, runs commands, and integrates with your development tools across terminal, IDE, desktop app, and browser.&lt;/p&gt;

&lt;p&gt;That matters because coding agents need somewhere to work.&lt;/p&gt;

&lt;p&gt;A Claude Code tutorial in 2026 is no longer just “install the CLI and ask it to fix a bug.” It has to cover permissions, IDE handoff, desktop sessions, hooks, MCP, CI workflows, and how to keep long-running work understandable.&lt;/p&gt;

&lt;p&gt;The June 2026 &lt;a href="https://claude.com/blog/artifacts-in-claude-code" rel="noopener noreferrer"&gt;Claude Code artifacts update&lt;/a&gt; makes that shift clearer. Artifacts turn work in progress into live, shareable pages: PR walkthroughs, system explainers, dashboards, release checklists, and incident pages built from the full session context.&lt;/p&gt;

&lt;p&gt;From a support-engineering angle, that is more valuable than another flashy benchmark. A coding agent that changes code but cannot explain the work creates operational debt. A coding agent that leaves behind a readable artifact gives the next human a fighting chance.&lt;/p&gt;

&lt;p&gt;The Alibaba angle also needs careful wording.&lt;/p&gt;

&lt;p&gt;Alibaba Cloud documents a way to connect Claude Code to Alibaba Cloud Model Studio through Anthropic-compatible endpoints and model mappings in its &lt;a href="https://www.alibabacloud.com/help/en/model-studio/claude-code" rel="noopener noreferrer"&gt;Claude Code Model Studio guide&lt;/a&gt;. Separately, Alibaba Cloud’s &lt;a href="https://www.alibabacloud.com/help/en/lingma/product-overview/introduction-of-lingma" rel="noopener noreferrer"&gt;Qoder CN suite&lt;/a&gt; is a family of domestic AI agent products for coding, work, terminal use, and cloud agents, built around Chinese models and China-based deployment.&lt;/p&gt;

&lt;p&gt;So I would not call this simply “Alibaba’s Claude Code.” That blurs the facts.&lt;/p&gt;

&lt;p&gt;The better read is this: Claude-Code-like workflows are becoming important enough that cloud vendors now want local model routing, domestic deployment, quota systems, and compliance-friendly agent environments.&lt;/p&gt;

&lt;p&gt;That is a much bigger story than a clone.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNDgxNWMwNTlkYWY4YmQzMDgxM2UxMDM2OWE2MmY2ZDJfeTByN0FZR0lnU2RMSWFDTzdSMlV1QUJJRVdJRDFEZE1fVG9rZW46S1dEOGJxQnowb0U2U1B4cWtSTWMzbkVUbkpnXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNDgxNWMwNTlkYWY4YmQzMDgxM2UxMDM2OWE2MmY2ZDJfeTByN0FZR0lnU2RMSWFDTzdSMlV1QUJJRVdJRDFEZE1fVG9rZW46S1dEOGJxQnowb0U2U1B4cWtSTWMzbkVUbkpnXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Claude Code Documentation Overview | Anthropic AI Agentic Tool" width="1149" height="628"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo from &lt;a href="https://code.claude.com/docs/en/overview" rel="noopener noreferrer"&gt;Claude&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cursor 2.0 Shows the IDE Is Becoming Agent-First
&lt;/h2&gt;

&lt;p&gt;Cursor 2.0 is useful because it says the quiet part out loud.&lt;/p&gt;

&lt;p&gt;In its &lt;a href="https://cursor.com/blog/2-0" rel="noopener noreferrer"&gt;Cursor 2.0 announcement&lt;/a&gt;, Cursor described the release around two major updates: Composer, its coding model, and a new interface for working with many agents in parallel. The &lt;a href="https://cursor.com/changelog/2-0" rel="noopener noreferrer"&gt;Cursor 2.0 changelog&lt;/a&gt; also describes multi-agent workflows, up to eight agents in parallel, isolated copies of the codebase, improved code review, and Browser becoming generally available.&lt;/p&gt;

&lt;p&gt;That is the important “Cursor 2.0 new features” story. Not one isolated browser feature, but an IDE reorganizing itself around agents.&lt;/p&gt;

&lt;p&gt;Traditional IDEs are file-first. Agentic IDEs are task-first.&lt;/p&gt;

&lt;p&gt;You ask for an outcome. Agents branch off. They work in isolated environments. You compare results. You review the diffs. You test the running app. You open files when you need depth, not because files are the only interface.&lt;/p&gt;

&lt;p&gt;The browser piece fits naturally here. Cursor says Browser for Agent became GA in 2.0 and can be embedded in-editor, with tools to select elements and forward DOM information to the agent. That matters especially for frontend work, where a patch can be syntactically correct and visually wrong.&lt;/p&gt;

&lt;p&gt;A browser-aware coding agent can inspect the app it just changed. That is the difference between “I edited the CSS” and “I checked the actual screen.”&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZTgzNTMyMThmZDk2MWEwMjBkZjM4MWM5MzM3NDBhNTFfMUZZb0JvallEU3hTZHY2b2Ryb0NqbEtyZUszNDhqbVNfVG9rZW46RXNPR2J0VU5Fb1FjMjh4T3VDS2NxMlR0bnpoXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DZTgzNTMyMThmZDk2MWEwMjBkZjM4MWM5MzM3NDBhNTFfMUZZb0JvallEU3hTZHY2b2Ryb0NqbEtyZUszNDhqbVNfVG9rZW46RXNPR2J0VU5Fb1FjMjh4T3VDS2NxMlR0bnpoXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Cursor Composer Benchmark | Coding Intelligence vs Speed Chart" width="828" height="466"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo from &lt;a href="https://cursor.com/blog/2-0" rel="noopener noreferrer"&gt;Cursor&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Chunk Sidecars Close the Loop Agents Keep Breaking
&lt;/h2&gt;

&lt;p&gt;CircleCI’s Chunk Sidecars may be the least flashy product in this comparison. They may also be the most practical.&lt;/p&gt;

&lt;p&gt;CircleCI describes &lt;a href="https://circleci.com/blog/chunk-sidecars/" rel="noopener noreferrer"&gt;Chunk Sidecars&lt;/a&gt; as lightweight, preconfigured environments that run alongside local or agent workflows and validate changes as they happen. The sidecar mirrors the project stack, detects test commands and build systems, runs scoped “microbuilds,” and feeds results back to the agent before the code reaches full CI.&lt;/p&gt;

&lt;p&gt;That solves a real problem.&lt;/p&gt;

&lt;p&gt;AI agents make the inner loop faster. They create more diffs, more branches, and more attempted fixes. But if validation still happens only after push, CI becomes the cleanup crew. The agent has already moved on, the context is colder, and the human has to reconstruct what happened.&lt;/p&gt;

&lt;p&gt;Chunk Sidecars move validation into the agent’s working loop.&lt;/p&gt;

&lt;p&gt;A coding agent can make a change, run a targeted check, read the failure, fix the issue, and repeat before polluting the shared pipeline. CircleCI’s follow-up on &lt;a href="https://circleci.com/blog/chunk-sidecar-agent-hooks/" rel="noopener noreferrer"&gt;agent hooks&lt;/a&gt; goes even further, showing how validation can trigger automatically at checkpoints in the agent workflow.&lt;/p&gt;

&lt;p&gt;This is what AI coding tools need more of.&lt;/p&gt;

&lt;p&gt;Not more confidence. More evidence.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNzE3Zjc0ZDhkY2I5ZjY0Nzk0NDk2MGMzNWUwYmJlNjVfY0RDS2Rsc2ZyYnlBRnJKSjN2OGFJNEFRRDJzbU1nMnVfVG9rZW46SlZ2SWJCMmROb1M1OEd4bXpvZGNFMGJpbnRnXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNzE3Zjc0ZDhkY2I5ZjY0Nzk0NDk2MGMzNWUwYmJlNjVfY0RDS2Rsc2ZyYnlBRnJKSjN2OGFJNEFRRDJzbU1nMnVfVG9rZW46SlZ2SWJCMmROb1M1OEd4bXpvZGNFMGJpbnRnXzE3ODQyNzk2NjY6MTc4NDI4MzI2Nl9WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="Chunk CLI Sidecar Sync Terminal | Local Coding Workspace Tool" width="840" height="478"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Photo from &lt;a href="https://circleci.com/blog/chunk-sidecars/" rel="noopener noreferrer"&gt;circleci Blog&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Would Actually Compare AI Coding Tools in 2026
&lt;/h2&gt;

&lt;p&gt;A serious AI coding tools comparison should not start with “which one writes the most impressive demo?”&lt;/p&gt;

&lt;p&gt;It should start with questions like these:&lt;/p&gt;

&lt;p&gt;Can the tool see the system it is editing?&lt;/p&gt;

&lt;p&gt;Can it run tests before CI?&lt;/p&gt;

&lt;p&gt;Can it explain its work to another human?&lt;/p&gt;

&lt;p&gt;Can I control permissions and cost?&lt;/p&gt;

&lt;p&gt;Can I compare multiple agent attempts without creating chaos?&lt;/p&gt;

&lt;p&gt;Can the workflow survive handoff, review, and rollback?&lt;/p&gt;

&lt;p&gt;Codex is pushing toward agent control, usage management, and physical workflow surfaces.&lt;/p&gt;

&lt;p&gt;Claude Code is pushing toward long-running, inspectable, cross-surface development work.&lt;/p&gt;

&lt;p&gt;Cursor 2.0 is pushing the IDE toward multi-agent, browser-aware coding.&lt;/p&gt;

&lt;p&gt;CircleCI Chunk Sidecars are pushing validation closer to the agent’s inner loop.&lt;/p&gt;

&lt;p&gt;Alibaba Cloud’s Model Studio and Qoder CN show the same pattern under local deployment, Chinese model routing, and compliance pressure.&lt;/p&gt;

&lt;p&gt;The products are different. The direction is shared.&lt;/p&gt;

&lt;p&gt;AI coding is becoming less like autocomplete and more like an operating layer for software delivery.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Shift
&lt;/h2&gt;

&lt;p&gt;The headline is not “Codex got a keyboard.”&lt;/p&gt;

&lt;p&gt;The headline is that coding agents now need keyboards, browser state, credit controls, artifacts, hooks, sidecars, model routing, and CI-adjacent validation.&lt;/p&gt;

&lt;p&gt;That is what happens when a tool moves from toy to infrastructure.&lt;/p&gt;

&lt;p&gt;When AI writes a snippet, a prompt is enough.&lt;/p&gt;

&lt;p&gt;When AI edits production code, you need approvals.&lt;/p&gt;

&lt;p&gt;When AI runs longer tasks, you need budget controls.&lt;/p&gt;

&lt;p&gt;When AI changes UI, you need a browser.&lt;/p&gt;

&lt;p&gt;When AI opens PRs, you need validation before CI.&lt;/p&gt;

&lt;p&gt;When AI touches a real team, you need artifacts and handoff.&lt;/p&gt;

&lt;p&gt;So the 2026 question is not just “Claude Code vs Cursor vs Codex.”&lt;/p&gt;

&lt;p&gt;It is: how complete does your development loop need to be?&lt;/p&gt;

&lt;p&gt;The best AI coding tool is not the one that writes the most code.&lt;/p&gt;

&lt;p&gt;It is the one that leaves the least mess after the code is written.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>openai</category>
      <category>coding</category>
    </item>
    <item>
      <title>LingBot-Video Isn’t Trying to Make Better Videos for Humans</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Fri, 10 Jul 2026 03:45:00 +0000</pubDate>
      <link>https://dev.to/ariakovac/lingbot-video-isnt-trying-to-make-better-videos-for-humans-3oo3</link>
      <guid>https://dev.to/ariakovac/lingbot-video-isnt-trying-to-make-better-videos-for-humans-3oo3</guid>
      <description>&lt;p&gt;A normal video model tries to make something a person wants to watch.&lt;/p&gt;

&lt;p&gt;A robot video model has a strange job. It needs to make something a robot can learn from.&lt;/p&gt;

&lt;p&gt;That difference sounds obvious until you start looking at what gets optimized. Human-facing models are usually judged by visual fidelity, style control, prompt following, camera motion, and whether the result feels coherent to us. Embodied video models need a different kind of coherence: object permanence, contact physics, action consequences, egocentric motion, manipulation intent, and whether the generated scene could actually teach a robot something useful.&lt;/p&gt;

&lt;p&gt;That is why the LingBot series is worth paying attention to.&lt;/p&gt;

&lt;p&gt;The newly released &lt;a href="https://arxiv.org/abs/2607.07675" rel="noopener noreferrer"&gt;LingBot-Video paper&lt;/a&gt; describes a DiT-based mixture-of-experts video foundation model built specifically for embodied intelligence. The authors call it an inaugural large-scale, open-source MoE video foundation model for robotics, which I would repeat carefully as the paper’s framing rather than as a settled industry label.&lt;/p&gt;

&lt;p&gt;Still, the direction is clear: video generation is moving from cinematic output toward robot training infrastructure.&lt;/p&gt;

&lt;p&gt;As someone who spends a lot of time thinking about messy workflows, this is the part I find interesting. The question is not whether LingBot-Video makes prettier clips than Sora-style models. The question is whether video generation can become part of the data loop for embodied AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Robot Video Is a Different Problem
&lt;/h2&gt;

&lt;p&gt;When a text-to-video model makes a cup slide strangely across a table, a human viewer may forgive it if the shot still looks beautiful.&lt;/p&gt;

&lt;p&gt;A robot cannot forgive that.&lt;/p&gt;

&lt;p&gt;If the cup moves without contact, the model is teaching the wrong physics. If a gripper passes through an object, the trajectory is useless. If the camera motion looks smooth but breaks spatial consistency, an embodied agent may learn a shortcut that fails in the real world.&lt;/p&gt;

&lt;p&gt;That is the difference between video as media and video as training signal.&lt;/p&gt;

&lt;p&gt;LingBot-Video addresses this by focusing its data and reward design around robot-relevant scenes. The paper describes a data profiling engine that augments generic internet video with embodied data sources: robot manipulation, robot navigation, and egocentric human video. It also introduces reward modeling around physical rationality and task completion, not just visual quality.&lt;/p&gt;

&lt;p&gt;That matters because robot training data is expensive.&lt;/p&gt;

&lt;p&gt;You can scrape enormous amounts of internet video, but most of it was not filmed from the perspective of an agent acting in the world. It rarely contains clean action labels. It may show objects, but not the causal structure a robot needs: what changed because of which action.&lt;/p&gt;

&lt;p&gt;So the value of a model like LingBot-Video is not simply that it generates videos. It is that it tries to generate videos in a distribution closer to embodied learning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why MoE Makes Sense Here
&lt;/h2&gt;

&lt;p&gt;The MoE part is not just a buzzword.&lt;/p&gt;

&lt;p&gt;Robotics is not one task. Navigation, grasping, tool use, object interaction, and egocentric prediction all stress different parts of a model. A dense model has to carry all of that capacity everywhere. A mixture-of-experts model can route tokens or conditions through different expert pathways.&lt;/p&gt;

&lt;p&gt;In practice, the promise is simple: more capacity without always paying the full compute cost.&lt;/p&gt;

&lt;p&gt;That is especially relevant for embodied AI, where the data distribution is broad and awkward. A kitchen manipulation clip is not the same as a corridor navigation clip. A third-person video is not the same as a head-mounted view. A model that treats them all as one flat video problem may miss the structure that matters.&lt;/p&gt;

&lt;p&gt;For developers, I would watch whether LingBot-Video’s expert routing actually improves downstream robotics tasks, not only whether the generated samples look convincing. If the model produces prettier robot videos but does not improve policy learning, it is still mostly a demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  LingBot-World 2.0 Is the Other Half of the Story
&lt;/h2&gt;

&lt;p&gt;LingBot-Video is about embodied video generation. LingBot-World 2.0 is closer to a world model.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://arxiv.org/abs/2607.07534" rel="noopener noreferrer"&gt;LingBot-World 2.0 paper&lt;/a&gt;, titled “Infinite Worlds with Versatile Interactions,” reports an interactive world model aimed at long-horizon generation and action-controllable simulation. The authors describe a pilot/director style generation harness and a distilled real-time variant targeting 720p video streams at 60fps.&lt;/p&gt;

&lt;p&gt;This distinction matters.&lt;/p&gt;

&lt;p&gt;A video foundation model can generate plausible embodied video. A world model tries to maintain an interactive environment over time. The second problem is harder because errors accumulate. A small spatial inconsistency at minute two can become nonsense by minute twenty.&lt;/p&gt;

&lt;p&gt;I would treat LingBot-World 2.0’s reported capabilities as research claims to evaluate, not production guarantees.&lt;/p&gt;

&lt;p&gt;But direction is important.&lt;/p&gt;

&lt;p&gt;If a robot can train inside a reasonably controllable generated world before touching the real world, the cost curve changes. Not because simulation replaces reality. It will not. But because it can make exploration cheaper before expensive real-world validation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where VLA Models Fit
&lt;/h2&gt;

&lt;p&gt;A VLA model is still a different layer.&lt;/p&gt;

&lt;p&gt;Vision-language-action models consume visual state and language goals, then produce actions. LingBot-VLA 2.0, described in a separate &lt;a href="https://arxiv.org/abs/2607.06403" rel="noopener noreferrer"&gt;technical report&lt;/a&gt;, focuses on embodied control using a large dataset that combines robot trajectories and egocentric human videos.&lt;/p&gt;

&lt;p&gt;So I would separate the stack like this:&lt;/p&gt;

&lt;p&gt;LingBot-Video generates embodied video distributions.&lt;/p&gt;

&lt;p&gt;LingBot-World 2.0 simulates longer interactive environments.&lt;/p&gt;

&lt;p&gt;LingBot-VLA 2.0 turns perception and language into robot actions.&lt;/p&gt;

&lt;p&gt;Those layers can support each other, but they are not interchangeable.&lt;/p&gt;

&lt;p&gt;This is also where I think a lot of quick coverage gets fuzzy. Calling everything a world model or everything a VLA model hides the engineering question: what part of the robot learning loop does this component actually improve?&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Test Before Getting Excited
&lt;/h2&gt;

&lt;p&gt;If I were evaluating LingBot-Video for a real robotics data loop, I would not start with the nicest demo video.&lt;/p&gt;

&lt;p&gt;I would ask five boring questions.&lt;/p&gt;

&lt;p&gt;First, does generated data improve downstream policy performance compared with real-only training?&lt;/p&gt;

&lt;p&gt;Second, does it help on tasks outside the generation distribution, or only on visually similar benchmarks?&lt;/p&gt;

&lt;p&gt;Third, does the model preserve contact physics well enough for manipulation?&lt;/p&gt;

&lt;p&gt;Fourth, can failures be filtered automatically, or does the pipeline still need heavy human review?&lt;/p&gt;

&lt;p&gt;Fifth, what happens when generated data is fed back into training repeatedly?&lt;/p&gt;

&lt;p&gt;That last one worries me the most.&lt;/p&gt;

&lt;p&gt;Synthetic robot video can be useful, but it can also amplify model mistakes. If a world model learns slightly wrong physics and a policy trains on that at scale, the policy may become very good at a world that does not exist.&lt;/p&gt;

&lt;p&gt;That is not a reason to ignore the approach. It is a reason to build verification into the loop from the start.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Shift: Video as Robot Infrastructure
&lt;/h2&gt;

&lt;p&gt;From Sora-style models to LingBot, the shift is not simply better video.&lt;/p&gt;

&lt;p&gt;It is a shift in purpose.&lt;/p&gt;

&lt;p&gt;For human viewers, video generation is a creative medium. For robots, video generation may become a training substrate: a way to create scenarios, stress-test policies, expand rare cases, and reduce the cost of physical trial and error.&lt;/p&gt;

&lt;p&gt;The next useful benchmark will not be whether the robot video looks cool on social media. It will be whether a robot trained with it becomes more reliable, more sample-efficient, and less surprised by the real world.&lt;/p&gt;

&lt;p&gt;As a developer, that is the line I care about.&lt;/p&gt;

&lt;p&gt;Not: did the model make a better video?&lt;/p&gt;

&lt;p&gt;But: did the generated world teach the agent something true?&lt;/p&gt;

&lt;p&gt;That is why LingBot-Video feels like a meaningful signal even if the field is still early.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Prompt Engineering Isn’t Dead. It Just Got Wrapped in a Harness.</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Thu, 09 Jul 2026 05:48:22 +0000</pubDate>
      <link>https://dev.to/ariakovac/prompt-engineering-isnt-dead-it-just-got-wrapped-in-a-harness-50c</link>
      <guid>https://dev.to/ariakovac/prompt-engineering-isnt-dead-it-just-got-wrapped-in-a-harness-50c</guid>
      <description>&lt;p&gt;A few months ago, if someone on my team said an AI workflow was “bad,” the first question was usually: &lt;em&gt;What prompt are we using?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Now I find myself asking a different question:&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Where does the model get feedback, and what is allowed to change after it fails?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That shift sounds small, but it changes almost everything.&lt;/p&gt;

&lt;p&gt;Lilian Weng’s July 2026 post, &lt;a href="https://lilianweng.github.io/posts/2026-07-04-harness/" rel="noopener noreferrer"&gt;“Harness Engineering for Self-Improvement”&lt;/a&gt;, helped frame a pattern many agent builders have already been wrestling with in practice. The model is still important. Of course it is. But the system wrapped around the model — tools, memory, workflow, permissions, evaluation, file state, background jobs, feedback loops — is becoming just as important as the prompt itself.&lt;/p&gt;

&lt;p&gt;Weng defines a harness as the system around a base model that orchestrates execution: how the model plans, calls tools, acts, manages context, stores artifacts, and evaluates results.&lt;/p&gt;

&lt;p&gt;That is the part that caught my attention.&lt;/p&gt;

&lt;p&gt;I work close enough to customer support workflows to know that “better AI” rarely means “one better answer.” In practice, it means you stop getting the same repeated failures, the system actually recovers instead of freezing when things get ambiguous, and human handoffs become a whole lot smoother. Plus, you finally get a system that doesn't completely forget why something broke last Tuesday.&lt;/p&gt;

&lt;p&gt;That is not prompt engineering anymore.&lt;/p&gt;

&lt;p&gt;That is harness design.&lt;/p&gt;

&lt;h2&gt;
  
  
  A Prompt Is an Instruction. A Harness Is a Working Environment.
&lt;/h2&gt;

&lt;p&gt;Prompt engineering is not useless. I still care about clear goals, constraints, examples, and tone.&lt;/p&gt;

&lt;p&gt;But a prompt is mostly instruction.&lt;/p&gt;

&lt;p&gt;A harness is environment.&lt;/p&gt;

&lt;p&gt;The difference matters because most agent failures are not just wording problems. They are system problems:&lt;/p&gt;

&lt;p&gt;The model gets too much context and loses the relevant bit.&lt;/p&gt;

&lt;p&gt;The tool result is noisy.&lt;/p&gt;

&lt;p&gt;The agent retries the same broken action.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMzkwODQwYzQzOTQ1NzIxYmM0OTg2NDk3YmYyY2RhY2RfUGlFemJkSUFRVHJMT3laTmVEYVc5aDBZZDVRNjJLYmZfVG9rZW46VUdJcGJJbFhtb1FNMkt4RlV6NWNRUGxLbk9oXzE3ODM1NzU4Njc6MTc4MzU3OTQ2N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DMzkwODQwYzQzOTQ1NzIxYmM0OTg2NDk3YmYyY2RhY2RfUGlFemJkSUFRVHJMT3laTmVEYVc5aDBZZDVRNjJLYmZfVG9rZW46VUdJcGJJbFhtb1FNMkt4RlV6NWNRUGxLbk9oXzE3ODM1NzU4Njc6MTc4MzU3OTQ2N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image about ai agent" width="1170" height="780"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The workflow has no memory of previous failures.&lt;/p&gt;

&lt;p&gt;The evaluation step is vague or missing.&lt;/p&gt;

&lt;p&gt;The system cannot tell whether an answer is merely plausible or actually verified.&lt;/p&gt;

&lt;p&gt;You can patch some of this with a better prompt. For a while.&lt;/p&gt;

&lt;p&gt;But eventually the prompt becomes a crowded apartment: instructions, warnings, examples, tool notes, edge cases, memory snippets, formatting rules, and “please don’t do that weird thing again” all squeezed into one place.&lt;/p&gt;

&lt;p&gt;A harness gives those responsibilities proper rooms.&lt;/p&gt;

&lt;p&gt;Memory can live in files or structured stores.&lt;/p&gt;

&lt;p&gt;Evaluation can run as a check, not a paragraph of advice.&lt;/p&gt;

&lt;p&gt;Tools can have permissions.&lt;/p&gt;

&lt;p&gt;Retries can have policy.&lt;/p&gt;

&lt;p&gt;Logs can become evidence.&lt;/p&gt;

&lt;p&gt;Failures can be mined instead of forgotten.&lt;/p&gt;

&lt;p&gt;That is why I do not read “harness vs prompt engineering” as a replacement story. It is more like a promotion. Prompt engineering becomes one component inside a larger runtime system.&lt;/p&gt;

&lt;h2&gt;
  
  
  Forget Code Rewriting: The Real Reality of AI Self-Improvement
&lt;/h2&gt;

&lt;p&gt;The hype around AI self-improvement loves to paint a dramatic picture: a model autonomously rewriting its own weights, training its next iteration, and scaling up recursively. But that’s the Hollywood version.&lt;/p&gt;

&lt;p&gt;Maybe someday. But near-term self-improvement looks much more boring and much more useful.&lt;/p&gt;

&lt;p&gt;It starts outside the model.&lt;/p&gt;

&lt;p&gt;In Weng’s framing, the practical near-term path is not necessarily a model directly editing its own weights. It is the harness becoming an optimization target. The system learns better ways to select context, manage state, run workflows, evaluate outputs, and preserve useful artifacts.&lt;/p&gt;

&lt;p&gt;That matches what I see in developer tooling.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYWYwNDJkOWJmMzdmNzUyOTgzYjZhNWViMDQ1ZjllZGVfck0zUW9JTTltN2YxcjR6NTJ2SkFOR2FGcUlhS0FLZzRfVG9rZW46UE43MWJMaFRVb1JvZEd4OWRGTWNWdXpibnNiXzE3ODM1NzU4Njc6MTc4MzU3OTQ2N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYWYwNDJkOWJmMzdmNzUyOTgzYjZhNWViMDQ1ZjllZGVfck0zUW9JTTltN2YxcjR6NTJ2SkFOR2FGcUlhS0FLZzRfVG9rZW46UE43MWJMaFRVb1JvZEd4OWRGTWNWdXpibnNiXzE3ODM1NzU4Njc6MTc4MzU3OTQ2N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="ai robots learns better ways to work" width="1332" height="749"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If an agent fails a coding task, maybe the model is not smart enough. But maybe the harness failed first:&lt;/p&gt;

&lt;p&gt;It did not inspect the right file.&lt;/p&gt;

&lt;p&gt;It did not run the failing test.&lt;/p&gt;

&lt;p&gt;It did not preserve the error trace.&lt;/p&gt;

&lt;p&gt;It did not know when to ask for clarification.&lt;/p&gt;

&lt;p&gt;It did not compare the final patch against the original requirement.&lt;/p&gt;

&lt;p&gt;This is where AI agent loop design becomes interesting. A minimal loop might look like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Plan the next action.&lt;/li&gt;
&lt;li&gt;Execute through a tool.&lt;/li&gt;
&lt;li&gt;Observe the result.&lt;/li&gt;
&lt;li&gt;Store useful evidence.&lt;/li&gt;
&lt;li&gt;Evaluate against the goal.&lt;/li&gt;
&lt;li&gt;Update the workflow or memory.&lt;/li&gt;
&lt;li&gt;Try again with better state.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That loop is not magic. It is just software engineering applied to agent behavior.&lt;/p&gt;

&lt;p&gt;But the moment you can actually look under the hood and tweak that loop, the system finds a real path to self-improvement—without having to pretend the model magically became its own R&amp;amp;D lab overnight.&lt;/p&gt;

&lt;h2&gt;
  
  
  The New Unit of Optimization Is the Run
&lt;/h2&gt;

&lt;p&gt;What I like about the harness framing is that it moves attention from the answer to the run.&lt;/p&gt;

&lt;p&gt;A run has traces. Tool calls. Files touched. Tests executed. Errors ignored. Assumptions made. Human interventions. Final evidence.&lt;/p&gt;

&lt;p&gt;That means the system can learn from more than success or failure.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://arxiv.org/abs/2606.09498" rel="noopener noreferrer"&gt;Self-Harness paper&lt;/a&gt; makes this explicit. It describes a loop where an LLM-based agent improves its own operating harness through weakness mining, harness proposal, and proposal validation. The important part is not “the agent rewrites itself” in some vague way. It is that failures are clustered, proposed harness changes are bounded, and accepted edits must pass regression checks.&lt;/p&gt;

&lt;p&gt;That last part matters.&lt;/p&gt;

&lt;p&gt;Self-improvement without regression testing is just automation with confidence issues.&lt;/p&gt;

&lt;p&gt;The same pattern appears in newer harness research around coding agents. &lt;a href="https://arxiv.org/abs/2604.25850" rel="noopener noreferrer"&gt;Agentic Harness Engineering&lt;/a&gt; frames harness evolution around observability: editable components, distilled experience, and decisions that can later be checked against outcomes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNzY1ZjUxYzlmZmZiNzFhYzAyZjE3ZWQ3ODI5OTBlYmZfbktPV0taRjhDMDN6dERtUTNJMlRnSVBZOE5BeGlJdmNfVG9rZW46QVlnemJ1ZHdlb0dtRER4UUljaWNFcU0ybkFlXzE3ODM1NzU4Njc6MTc4MzU3OTQ2N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DNzY1ZjUxYzlmZmZiNzFhYzAyZjE3ZWQ3ODI5OTBlYmZfbktPV0taRjhDMDN6dERtUTNJMlRnSVBZOE5BeGlJdmNfVG9rZW46QVlnemJ1ZHdlb0dtRER4UUljaWNFcU0ybkFlXzE3ODM1NzU4Njc6MTc4MzU3OTQ2N19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="1158" height="580"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is the version of AI self-improvement I can actually imagine trusting in production.&lt;/p&gt;

&lt;p&gt;Not an agent randomly rewriting its system prompt because it “reflected.”&lt;/p&gt;

&lt;p&gt;A system that says:&lt;/p&gt;

&lt;p&gt;“This failure pattern happened 14 times.”&lt;/p&gt;

&lt;p&gt;“The root cause appears to be missing project-state recovery.”&lt;/p&gt;

&lt;p&gt;“Here is a narrow harness change.”&lt;/p&gt;

&lt;p&gt;“Here is the regression set.”&lt;/p&gt;

&lt;p&gt;“Here is what improved, and here is what did not.”&lt;/p&gt;

&lt;p&gt;That is less cinematic. It is also much closer to how reliable engineering works.&lt;/p&gt;

&lt;h2&gt;
  
  
  ECC Shows the Developer Version of This Shift
&lt;/h2&gt;

&lt;p&gt;The GitHub repo &lt;a href="https://github.com/affaan-m/ECC" rel="noopener noreferrer"&gt;&lt;code&gt;affaan-m/ECC&lt;/code&gt;&lt;/a&gt; is a useful signal because it turns the harness idea into something developers can poke at.&lt;/p&gt;

&lt;p&gt;The repo describes itself as a harness-native operator system for agentic work, spanning skills, memory optimization, security scanning, cross-harness workflows, and support for tools like Claude Code, Codex, Cursor, OpenCode, and others.&lt;/p&gt;

&lt;p&gt;I would not treat any GitHub project as proof that a concept is mature. But I do think ECC shows why the harness idea is spreading.&lt;/p&gt;

&lt;p&gt;Developers do not just want a smarter model.&lt;/p&gt;

&lt;p&gt;They want agents that remember project conventions.&lt;/p&gt;

&lt;p&gt;They want repeatable workflows.&lt;/p&gt;

&lt;p&gt;They want security boundaries.&lt;/p&gt;

&lt;p&gt;They want skills that can transfer across tools.&lt;/p&gt;

&lt;p&gt;They want logs, hooks, rules, and recovery paths.&lt;/p&gt;

&lt;p&gt;That is exactly the layer prompt engineering does not fully cover.&lt;/p&gt;

&lt;p&gt;A prompt can tell an agent to be careful.&lt;/p&gt;

&lt;p&gt;A harness can prevent it from touching production secrets, require tests before completion, preserve failure traces, and route uncertain cases to a human.&lt;/p&gt;

&lt;p&gt;That difference is not cosmetic. It is the difference between advice and infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  So Is Prompt Engineering Over?
&lt;/h2&gt;

&lt;p&gt;No.&lt;/p&gt;

&lt;p&gt;But the era where prompt engineering was the main interface for improving AI systems is fading.&lt;/p&gt;

&lt;p&gt;A good prompt still matters. It tells the model what the work is. It shapes behavior. It encodes constraints.&lt;/p&gt;

&lt;p&gt;But for long-running agents, the harder problems now sit around the prompt:&lt;/p&gt;

&lt;p&gt;What context should be loaded?&lt;/p&gt;

&lt;p&gt;Which tool should be available?&lt;/p&gt;

&lt;p&gt;What should be persisted?&lt;/p&gt;

&lt;p&gt;What counts as done?&lt;/p&gt;

&lt;p&gt;Who verifies the result?&lt;/p&gt;

&lt;p&gt;What happens after failure?&lt;/p&gt;

&lt;p&gt;What is allowed to change next time?&lt;/p&gt;

&lt;p&gt;That is why “AI Harness” feels like a more useful keyword than another prompt trick.&lt;/p&gt;

&lt;p&gt;For a small team, the practical move is not to build a grand self-improving agent on day one. It is to make your current loop visible.&lt;/p&gt;

&lt;p&gt;Start with one workflow.&lt;/p&gt;

&lt;p&gt;Log every tool call.&lt;/p&gt;

&lt;p&gt;Save the important artifacts.&lt;/p&gt;

&lt;p&gt;Define a verifier.&lt;/p&gt;

&lt;p&gt;Separate memory from instructions.&lt;/p&gt;

&lt;p&gt;Let failures produce structured notes.&lt;/p&gt;

&lt;p&gt;Only then consider letting the system propose changes to the harness.&lt;/p&gt;

&lt;p&gt;The scaling law bottleneck is not that models stopped improving. It is that better base models do not automatically give you better deployed systems. The gap between capability and reliability is increasingly filled by harnesses.&lt;/p&gt;

&lt;p&gt;That is where I think Lilian Weng’s post lands hardest.&lt;/p&gt;

&lt;p&gt;Prompt engineering asks: &lt;em&gt;What should I say to the model?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Harness engineering is really about one big question: how do you build a system that lets a model screw up, learn from it, verify the fix, and try again—all without turning your entire workflow into a total black box of hidden states? Having watched countless shiny AI features blow up in actual user-facing apps, I know exactly which problem I’d rather spend my time solving.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>The AI Race Is Becoming a Routing Problem, Not a Size Contest</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Mon, 06 Jul 2026 07:00:00 +0000</pubDate>
      <link>https://dev.to/ariakovac/the-ai-race-is-becoming-a-routing-problem-not-a-size-contest-ij7</link>
      <guid>https://dev.to/ariakovac/the-ai-race-is-becoming-a-routing-problem-not-a-size-contest-ij7</guid>
      <description>&lt;p&gt;Last week, I was looking at a support dashboard with three very unglamorous columns: average response latency, failed requests, and cost per resolved ticket.&lt;/p&gt;

&lt;p&gt;No launch video. No leaderboard screenshot. No beautiful demo where everything works on the first try.&lt;/p&gt;

&lt;p&gt;But that is usually where AI becomes real for me.&lt;/p&gt;

&lt;p&gt;I work close enough to customer support workflows to know that “state-of-the-art” only matters if it survives messy inputs, multilingual users, peak-hour traffic, and a finance team asking why the API bill doubled. So when I look at the current wave of AI infrastructure news — Meituan’s LongCat line, Google’s Nano Banana 2, world-model research like World-VLA-Loop, and Baidu’s &lt;code&gt;Unlimited-OCR&lt;/code&gt; — I do not read them as separate stories.&lt;/p&gt;

&lt;p&gt;I read them as one story: AI is moving from “how powerful can we make this model?” to “how much useful capability can we deliver per dollar, per second, and per deployment target?”&lt;/p&gt;

&lt;h2&gt;
  
  
  LongCat-2.0 Is Interesting, but the Safer Signal Is LongCat’s Efficiency Pattern
&lt;/h2&gt;

&lt;p&gt;The most sensitive LongCat headline right now is LongCat-2.0. A recent AFP report, picked up by Omni, says Meituan has launched LongCat-2.0 and that the company claims it was trained only on domestically made chips in its size class. That is an important claim, but I would treat it carefully until more technical details are publicly available.&lt;/p&gt;

&lt;p&gt;The better-documented technical signal is Meituan’s earlier LongCat work.&lt;/p&gt;

&lt;p&gt;In the &lt;a href="https://arxiv.org/abs/2509.01322" rel="noopener noreferrer"&gt;LongCat-Flash technical report&lt;/a&gt;, Meituan describes a 560B-parameter MoE language model that activates only 18.6B to 31.3B parameters per token, around 27B on average. The paper also reports more than 100 tokens per second for inference and a cost of \$0.70 per million output tokens.&lt;/p&gt;

&lt;p&gt;That is the part I care about as an engineer. The number on the box is huge, but the serving logic is about dynamic compute allocation.&lt;/p&gt;

&lt;p&gt;LongCat-Next pushes the same direction from another angle. The &lt;a href="https://arxiv.org/abs/2603.27538" rel="noopener noreferrer"&gt;LongCat-Next paper&lt;/a&gt; describes a native multimodal model that processes text, vision, and audio under one autoregressive objective.&lt;/p&gt;

&lt;p&gt;So for me, a real Meituan LongCat-2.0 review should not start with national tech excitement. It should ask boring production questions:&lt;/p&gt;

&lt;p&gt;Can p95 latency stay stable?&lt;/p&gt;

&lt;p&gt;What is the cost per million output tokens?&lt;/p&gt;

&lt;p&gt;How does the model behave on agentic tool-use tasks?&lt;/p&gt;

&lt;p&gt;Can the training and serving stack be repeated outside the usual Nvidia-centered assumptions?&lt;/p&gt;

&lt;p&gt;“Domestic chips for large-scale AI models” is not only a policy headline. For developers, it is a deployment question.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nano Banana 2 Makes Image Generation Feel More Operational
&lt;/h2&gt;

&lt;p&gt;Google’s Nano Banana 2 represents the other side of the efficiency race.&lt;/p&gt;

&lt;p&gt;Here, the story is not domestic training infrastructure. It is high-volume image generation becoming easier to price and route.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYzZjNDA3ZjE1NGY5YTNiNmQ5YjEzYWI1ZDEyMDUzY2ZfS29FbVhNQ0s3T3RzTGxZUGswVFNvUndlTGdtNVBKOEFfVG9rZW46VHlLc2JUSjV5b3RoUFV4aFhOcmM0alhObjJlXzE3ODMzMjAwNzM6MTc4MzMyMzY3M19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYzZjNDA3ZjE1NGY5YTNiNmQ5YjEzYWI1ZDEyMDUzY2ZfS29FbVhNQ0s3T3RzTGxZUGswVFNvUndlTGdtNVBKOEFfVG9rZW46VHlLc2JUSjV5b3RoUFV4aFhOcmM0alhObjJlXzE3ODMzMjAwNzM6MTc4MzMyMzY3M19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="760" height="482"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;According to the official &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Gemini API pricing page&lt;/a&gt;, Nano Banana 2 is a &lt;code&gt;gemini-3.1-flash-image&lt;/code&gt;. Standard pricing lists 1K image output at \$0.067 per image, while Batch pricing lists 1K image output at \$0.034 per image. The cheaper Lite model, &lt;code&gt;gemini-3.1-flash-lite-image&lt;/code&gt;, is listed at \$0.0336 per 1K image on standard pricing and \$0.0168 in Batch.&lt;/p&gt;

&lt;p&gt;That distinction matters. “\$0.034 per image” is real, but it is not automatically the price of every real-time user interaction. It depends on Batch mode, image size, retries, and whether the user expects instant feedback.&lt;/p&gt;

&lt;p&gt;Google’s &lt;a href="https://ai.google.dev/gemini-api/docs/image-generation" rel="noopener noreferrer"&gt;image generation docs&lt;/a&gt; position Nano Banana 2 as the general-purpose image model in the current Gemini image family, with support for 4K output, text rendering, world knowledge, grounding, and multiple reference-image workflows.&lt;/p&gt;

&lt;p&gt;For a small product team, I would split the decision like this:&lt;/p&gt;

&lt;p&gt;Use Nano Banana 2 Lite when speed and cost matter more than complex editing.&lt;/p&gt;

&lt;p&gt;Use Nano Banana 2 when you need better text rendering, higher-resolution output, grounding, or more reliable multi-reference work.&lt;/p&gt;

&lt;p&gt;Use a more expensive model only when the creative task actually justifies it.&lt;/p&gt;

&lt;p&gt;The benchmark conversation should not stop at “cheap API.” Product cost also includes retries, moderation, queueing, storage, and human review. A cheap image model can still become expensive if every third output needs to be regenerated.&lt;/p&gt;

&lt;h2&gt;
  
  
  World-VLA-Loop Shows Why “World Model” Still Means Research First
&lt;/h2&gt;

&lt;p&gt;The original brief mentioned Loop world-model research, so I checked the public sources carefully. The verifiable paper I found is &lt;a href="https://huggingface.co/papers/2602.06508" rel="noopener noreferrer"&gt;World-VLA-Loop&lt;/a&gt;, listed on Hugging Face Papers and published on arXiv in February 2026.&lt;/p&gt;

&lt;p&gt;This is not something I would describe as a ready-to-use product. It is research around a closed-loop framework connecting video world models and Vision-Language-Action policies for robotics.&lt;/p&gt;

&lt;p&gt;The project page explains the idea more concretely: train a world model, let a VLA policy roll out inside that world model, use failures to refine the system, then deploy and iterate. The &lt;a href="https://showlab.github.io/World-VLA-Loop/" rel="noopener noreferrer"&gt;World-VLA-Loop project page&lt;/a&gt; describes this as a cycle for improving both the world model and the policy.&lt;/p&gt;

&lt;p&gt;That is exciting, but I would keep the language cautious.&lt;/p&gt;

&lt;p&gt;World models matter because they try to preserve environment state over time. Objects should stay consistent. Actions should follow physical constraints. Failure trajectories should teach the system something useful.&lt;/p&gt;

&lt;p&gt;But from a production perspective, I would still ask:&lt;/p&gt;

&lt;p&gt;How quickly do errors accumulate?&lt;/p&gt;

&lt;p&gt;Does the model preserve object identity after viewpoint changes?&lt;/p&gt;

&lt;p&gt;Can developers inspect why a rollout failed?&lt;/p&gt;

&lt;p&gt;Does this work beyond controlled robotics benchmarks?&lt;/p&gt;

&lt;p&gt;This is where world models and support systems have something in common: the first response matters, but the fifth interaction is where the system usually reveals itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  Unlimited-OCR Is the Kind of AI Project Developers Actually Use
&lt;/h2&gt;

&lt;p&gt;The GitHub attention around &lt;a href="https://github.com/baidu/Unlimited-OCR" rel="noopener noreferrer"&gt;&lt;code&gt;baidu/Unlimited-OCR&lt;/code&gt;&lt;/a&gt; feels especially practical.&lt;/p&gt;

&lt;p&gt;OCR is not glamorous. It is invoices, scanned PDFs, screenshots, forms, contracts, tables, weird fonts, and long documents someone uploads right before close of business.&lt;/p&gt;

&lt;p&gt;That is why the repo is interesting. Its own README describes “one-shot long-horizon parsing,” shows Transformers inference, and includes deployment paths through vLLM and SGLang. It also includes multi-page and PDF parsing examples.&lt;/p&gt;

&lt;p&gt;This is the practical side of lightweight AI. Not one universal model, but specialized models that solve painful tasks cleanly.&lt;/p&gt;

&lt;p&gt;In a real stack, I would rather route OCR to a document parser, simple image generation to a low-cost image model, sensitive flows to local or private deployments, and complex reasoning to a stronger model only when needed.&lt;/p&gt;

&lt;p&gt;That sounds less exciting than “one model to rule them all,” but it is usually how reliable systems are built.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYjk3NTdhMWM2MTdjYjgxNTVmMWVjMzY2NWEzOTBhOWJfOVpGSHVOYzR5Q2JUZFZHVWtTYXJ5VVdRWFN5STZFN21fVG9rZW46RjlxNWJhdmc0b3pUQld4b28yamNRQTE3bk1mXzE3ODMzMjAwNzM6MTc4MzMyMzY3M19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fecngm1r8vhmv.feishu.cn%2Fspace%2Fapi%2Fbox%2Fstream%2Fdownload%2Fasynccode%2F%3Fcode%3DYjk3NTdhMWM2MTdjYjgxNTVmMWVjMzY2NWEzOTBhOWJfOVpGSHVOYzR5Q2JUZFZHVWtTYXJ5VVdRWFN5STZFN21fVG9rZW46RjlxNWJhdmc0b3pUQld4b28yamNRQTE3bk1mXzE3ODMzMjAwNzM6MTc4MzMyMzY3M19WNA%26add_watermark%3Dtrue%26scene_type%3DCCM" alt="image" width="1170" height="780"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The Real Trend Is Model Routing
&lt;/h2&gt;

&lt;p&gt;The common thread across LongCat, Nano Banana 2, World-VLA-Loop, and Unlimited-OCR is not that they all do the same thing.&lt;/p&gt;

&lt;p&gt;They obviously do not.&lt;/p&gt;

&lt;p&gt;The common thread is efficiency pressure.&lt;/p&gt;

&lt;p&gt;High performance still matters. But low cost, low latency, deployment flexibility, and task-specific reliability now matter just as much.&lt;/p&gt;

&lt;p&gt;For developers, the winning AI architecture in 2026 may look less like choosing one “best” model and more like building a routing layer:&lt;/p&gt;

&lt;p&gt;A powerful model for complex reasoning.&lt;/p&gt;

&lt;p&gt;A cheap image model for high-volume creative drafts.&lt;/p&gt;

&lt;p&gt;A specialized OCR model for document parsing.&lt;/p&gt;

&lt;p&gt;A local or edge model for privacy-sensitive workflows.&lt;/p&gt;

&lt;p&gt;A world model only where persistent simulation actually matters.&lt;/p&gt;

&lt;p&gt;That is the part of the AI race I find more useful than leaderboard screenshots.&lt;/p&gt;

&lt;p&gt;The future is not just bigger models. It is better decisions about when not to use the biggest model.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>nanobanana</category>
      <category>devops</category>
    </item>
    <item>
      <title>The Claude Code “Trojan” Panic Is Really About Trust</title>
      <dc:creator>Aria Kovac</dc:creator>
      <pubDate>Thu, 02 Jul 2026 09:02:26 +0000</pubDate>
      <link>https://dev.to/ariakovac/the-claude-code-trojan-panic-is-really-about-trust-2k9f</link>
      <guid>https://dev.to/ariakovac/the-claude-code-trojan-panic-is-really-about-trust-2k9f</guid>
      <description>&lt;p&gt;What I changed in my AI tooling audit after the latest Anthropic controversy&lt;/p&gt;

&lt;p&gt;Last week, a thread about Claude Code started moving through my feeds in three languages.&lt;/p&gt;

&lt;p&gt;In English, people called it geofencing.&lt;/p&gt;

&lt;p&gt;In Chinese, people called it a “Trojan.”&lt;/p&gt;

&lt;p&gt;In my work Slack, someone asked the practical question: “So, should we uninstall AI coding tools from company laptops?”&lt;/p&gt;

&lt;p&gt;That was the moment I paid attention.&lt;/p&gt;

&lt;p&gt;I work as a customer support engineer at a cross-border e-commerce company in Amsterdam. Most days, I deal with messy human systems: angry customers, multilingual tickets, half-broken automations, API logs, and the strange middle layer where software decisions become customer pain.&lt;/p&gt;

&lt;p&gt;So when developers argue about whether Claude Code secretly checked IP addresses, system language, proxy behavior, or account identity signals, I do not read it only as an AI news story.&lt;/p&gt;

&lt;p&gt;I read it as a trust story.&lt;/p&gt;

&lt;p&gt;And trust, once it enters a workflow, needs logging.&lt;/p&gt;

&lt;h2&gt;
  
  
  First: I Don’t Like the Word “Trojan” Here
&lt;/h2&gt;

&lt;p&gt;The word “Trojan” does a lot of emotional work.&lt;/p&gt;

&lt;p&gt;It suggests malware. It suggests deception. It suggests that a tool entered your system as one thing and behaved as another.&lt;/p&gt;

&lt;p&gt;Maybe that is why the term spread so fast. It captured the feeling many developers had: “Wait, this coding assistant I invited into my terminal may also be evaluating whether I am allowed to exist as a user?”&lt;/p&gt;

&lt;p&gt;That feeling is real.&lt;/p&gt;

&lt;p&gt;But I would still separate three things:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1. Malware behavior
2. Telemetry and fraud detection
3. Policy enforcement based on location or identity
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They are not the same.&lt;/p&gt;

&lt;p&gt;A tool that steals secrets is one category. A tool that collects usage data is another. A tool that enforces regional restrictions using IP, identity verification, device signals, or proxy detection is another again.&lt;/p&gt;

&lt;p&gt;The trust problem begins when users cannot tell which category they are in.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Important Part Is Not China
&lt;/h2&gt;

&lt;p&gt;A lot of the current discussion focuses on China.&lt;/p&gt;

&lt;p&gt;That makes sense. Public reporting has described Anthropic’s strict access limits for users in China, workarounds involving VPNs and relay services, and anti-proxy detection systems used to disrupt unauthorized access. WIRED also reported that Anthropic does not offer commercial Claude access in China or to Chinese-owned subsidiaries outside the country.&lt;/p&gt;

&lt;p&gt;But if you only read this as a China story, you miss the part that will affect everyone.&lt;/p&gt;

&lt;p&gt;The real question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What signals can an AI development tool collect from your local environment, and what decisions can the vendor make with those signals?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Today the signal might be region.&lt;/p&gt;

&lt;p&gt;Tomorrow it might be enterprise policy.&lt;/p&gt;

&lt;p&gt;Next month it might be “suspicious automation.”&lt;/p&gt;

&lt;p&gt;Later it might be whether your company, country, payment method, IDE, proxy, or usage pattern fits a risk model you cannot inspect.&lt;/p&gt;

&lt;p&gt;That is not science fiction. That is normal platform enforcement, arriving inside developer tools.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffx5d3xj5je4ka8co4pae.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ffx5d3xj5je4ka8co4pae.png" alt="Third-party reporting shows the geolocation issue is part of a larger access-control problem." width="800" height="431"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  My Small Audit Checklist
&lt;/h2&gt;

&lt;p&gt;After this story, I made a checklist for every AI tool that touches my work machine.&lt;/p&gt;

&lt;p&gt;Not because I think every vendor is malicious. I do not.&lt;/p&gt;

&lt;p&gt;Because “I trust this company” is not an audit control. It is a mood.&lt;/p&gt;

&lt;p&gt;Here is the first version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI Tool Audit

1. What local data can it read?
   - current working directory
   - file contents
   - shell history
   - git metadata
   - environment variables
   - system language / locale
   - OS and device metadata

2. What network calls does it make?
   - API endpoint
   - telemetry endpoint
   - update endpoint
   - crash reporting endpoint

3. What identity signals are linked?
   - account email
   - payment country
   - phone number
   - ID verification
   - organization domain
   - IP address / proxy signals

4. What actions can it perform?
   - edit files
   - run shell commands
   - install packages
   - commit code
   - call external tools

5. What happens if access is revoked?
   - can work continue?
   - are local files safe?
   - are logs available?
   - is there an export path?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is not paranoia. This is the same way I think about customer support automations.&lt;/p&gt;

&lt;p&gt;If a tool can touch the customer, it needs controls.&lt;/p&gt;

&lt;p&gt;If a tool can touch code, it needs better controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Local Environment Is Customer Data Too
&lt;/h2&gt;

&lt;p&gt;Support engineers learn this lesson early: metadata is not “nothing.”&lt;/p&gt;

&lt;p&gt;A customer’s language, country, refund history, device type, order timing, and support channel can reveal more than the message itself.&lt;/p&gt;

&lt;p&gt;Developers sometimes treat machine metadata as less sensitive because it feels technical.&lt;/p&gt;

&lt;p&gt;It is not.&lt;/p&gt;

&lt;p&gt;Your system language may reveal location or working context. Your IP may reveal office routing. Your project path may reveal a client name. Your git remote may reveal private infrastructure. Your environment variables may reveal things nobody should paste into a model, ever.&lt;/p&gt;

&lt;p&gt;So when an AI coding tool runs locally, the question is not only “does it send my source code?”&lt;/p&gt;

&lt;p&gt;The question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What else does it observe while helping me?&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Expect From Vendors
&lt;/h2&gt;

&lt;p&gt;I do not think AI companies should pretend regional compliance does not exist.&lt;/p&gt;

&lt;p&gt;Export controls, sanctions, enterprise agreements, abuse prevention, fraud detection, and model safety policies are real constraints. A vendor may have legal reasons to block access. I understand that.&lt;/p&gt;

&lt;p&gt;But I expect four things.&lt;/p&gt;

&lt;p&gt;First, say what signals are collected.&lt;/p&gt;

&lt;p&gt;Not in a 9,000-word privacy policy written for lawyers. In a developer-readable table.&lt;/p&gt;

&lt;p&gt;Second, say which decisions those signals influence.&lt;/p&gt;

&lt;p&gt;There is a difference between “we collect locale for UI language” and “we use locale as one signal in account enforcement.”&lt;/p&gt;

&lt;p&gt;Third, provide an audit mode.&lt;/p&gt;

&lt;p&gt;Let me run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ai-tool doctor --privacy
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and see what endpoints, config files, local permissions, and telemetry settings are active.&lt;/p&gt;

&lt;p&gt;Fourth, make rollback visible.&lt;/p&gt;

&lt;p&gt;If a controversial enforcement mechanism changes, publish the version, the behavior, and the migration path. Do not make developers reverse-engineer trust from rumor.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Changed Personally
&lt;/h2&gt;

&lt;p&gt;I did not uninstall every AI coding tool.&lt;/p&gt;

&lt;p&gt;That would be dramatic and not very useful.&lt;/p&gt;

&lt;p&gt;I changed how I use them.&lt;/p&gt;

&lt;p&gt;For personal projects, I still use AI tools freely, but I keep secrets out of the working directory. For work-adjacent experiments, I use separate folders, separate tokens, and less ambient access. For anything involving customers, I do not let an AI agent roam.&lt;/p&gt;

&lt;p&gt;My current rule is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;If I would not paste it into a support ticket,
I do not let an AI tool inspect it by default.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That rule is imperfect. It is also easy to remember, which means I actually follow it.&lt;/p&gt;

&lt;p&gt;I have a similar rule for my restaurant database, oddly enough. If a field will later be used for filtering, ranking, or automation, I define it clearly at the start. Otherwise future-me will build queries on vibes and regret it in Lisbon over bad clams.&lt;/p&gt;

&lt;p&gt;Same lesson. Different table.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Account Ban Problem
&lt;/h2&gt;

&lt;p&gt;The SEO-friendly version of this article would probably be “What to Do If Your Claude Account Gets Banned.”&lt;/p&gt;

&lt;p&gt;The honest version is less exciting:&lt;/p&gt;

&lt;p&gt;Do not build a critical workflow around a single account you do not control.&lt;/p&gt;

&lt;p&gt;Have an export path. Keep local copies of prompts and configs. Know which projects depend on which AI vendor. Keep a second model good enough for emergencies. Do not route sensitive work through random relay services because they are cheaper.&lt;/p&gt;

&lt;p&gt;Especially do not send company code or customer data through unofficial proxy tools.&lt;/p&gt;

&lt;p&gt;That is not a clever workaround. That is a data incident waiting politely in line.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Lesson
&lt;/h2&gt;

&lt;p&gt;The Claude Code controversy is not only about Anthropic.&lt;/p&gt;

&lt;p&gt;It is about the next phase of AI tooling.&lt;/p&gt;

&lt;p&gt;We are moving from chatbots we visit to agents we install. They sit closer to our files, terminals, repos, tickets, and internal systems. That makes them more useful.&lt;/p&gt;

&lt;p&gt;It also makes vague trust much more expensive.&lt;/p&gt;

&lt;p&gt;So my takeaway is not “Claude Code is safe” or “Claude Code is unsafe.”&lt;/p&gt;

&lt;p&gt;My takeaway is this:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Any AI tool powerful enough to help with real work is powerful enough to deserve an audit.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not a dramatic one. A practical one.&lt;/p&gt;

&lt;p&gt;What can it read? What can it send? What can it change? What can get my account blocked? What happens if it disappears tomorrow?&lt;/p&gt;

&lt;p&gt;If the answer is “I don’t know,” that is not a reason to panic.&lt;/p&gt;

&lt;p&gt;It is a reason to open the settings, read the docs, check the network calls, and write the first version of your own checklist.&lt;/p&gt;

&lt;p&gt;That is where trust starts.&lt;/p&gt;

&lt;p&gt;Not with a statement from a vendor.&lt;/p&gt;

&lt;p&gt;With a workflow you can inspect.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>security</category>
      <category>discuss</category>
    </item>
  </channel>
</rss>
