<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mehrad_Arbab</title>
    <description>The latest articles on DEV Community by Mehrad_Arbab (@mehradarbab).</description>
    <link>https://dev.to/mehradarbab</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4040036%2Fa6a2c740-1863-4f12-9368-8c9ea2175d13.jpeg</url>
      <title>DEV Community: Mehrad_Arbab</title>
      <link>https://dev.to/mehradarbab</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mehradarbab"/>
    <language>en</language>
    <item>
      <title>AI Agents Don't Need More Autonomy. They Need Proof They Finished the Job.</title>
      <dc:creator>Mehrad_Arbab</dc:creator>
      <pubDate>Thu, 08 Oct 2026 17:56:45 +0000</pubDate>
      <link>https://dev.to/mehradarbab/ai-agents-dont-need-more-autonomy-they-need-proof-they-finished-the-job-epe</link>
      <guid>https://dev.to/mehradarbab/ai-agents-dont-need-more-autonomy-they-need-proof-they-finished-the-job-epe</guid>
      <description>&lt;p&gt;I'm 19. I've built the foundation of an AI-native workspace called LocalDesk.&lt;/p&gt;

&lt;p&gt;And the deeper I go into building AI agents, the more convinced I become that we're asking the wrong question.&lt;/p&gt;

&lt;p&gt;Everyone wants more autonomous agents.&lt;/p&gt;

&lt;p&gt;Agents that code. Agents that browse. Agents that execute workflows. Agents that operate entire businesses.&lt;/p&gt;

&lt;p&gt;But there's a much more important question:&lt;/p&gt;

&lt;p&gt;When an AI agent tells you it's done, how do you know it's telling the truth?&lt;/p&gt;

&lt;p&gt;Not whether the model is lying intentionally.&lt;/p&gt;

&lt;p&gt;Whether the work actually happened.&lt;/p&gt;

&lt;p&gt;The "Done" Problem&lt;/p&gt;

&lt;p&gt;Imagine asking an AI agent to deploy an application.&lt;/p&gt;

&lt;p&gt;It modifies your code, pushes a commit, calls the deployment API, and responds:&lt;/p&gt;

&lt;p&gt;"Successfully deployed."&lt;/p&gt;

&lt;p&gt;Sounds great.&lt;/p&gt;

&lt;p&gt;Except the deployment failed.&lt;/p&gt;

&lt;p&gt;Or the wrong environment was updated.&lt;/p&gt;

&lt;p&gt;Or the application started successfully but the authentication flow broke.&lt;/p&gt;

&lt;p&gt;The agent completed a sequence of actions.&lt;/p&gt;

&lt;p&gt;But it didn't necessarily complete your objective.&lt;/p&gt;

&lt;p&gt;That distinction matters more as agents receive greater permissions.&lt;/p&gt;

&lt;p&gt;More Intelligence Isn't the Whole Solution&lt;/p&gt;

&lt;p&gt;We have increasingly capable language models.&lt;/p&gt;

&lt;p&gt;They can reason through problems, generate software, and coordinate tools.&lt;/p&gt;

&lt;p&gt;But reliability in the real world requires more than generating the next intelligent action.&lt;/p&gt;

&lt;p&gt;I think we need a clearer separation between four concepts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;Intent — What does the human actually want?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Execution — What actions did the agent perform?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Verification — What independent evidence shows the objective was achieved?&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;Recovery — What happens when reality doesn't match the plan?&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Without these distinctions, an agent can look incredibly capable while quietly creating operational risk.&lt;/p&gt;

&lt;p&gt;What We've Built at LocalDesk&lt;/p&gt;

&lt;p&gt;This is one of the architectural problems that shaped LocalDesk.&lt;/p&gt;

&lt;p&gt;We've built the core architecture of an AI-native workspace around three connected systems.&lt;/p&gt;

&lt;p&gt;Baymax — Human Intent&lt;/p&gt;

&lt;p&gt;Baymax is our conversational AI interface.&lt;/p&gt;

&lt;p&gt;The goal is to let humans describe what they want to accomplish without manually coordinating every application and workflow.&lt;/p&gt;

&lt;p&gt;Missions — Execution and Verification&lt;/p&gt;

&lt;p&gt;Missions is our orchestration layer.&lt;/p&gt;

&lt;p&gt;Our architecture separates routing, context, planning, execution, verification, and repair.&lt;/p&gt;

&lt;p&gt;Instead of treating a model's response as the final result, we've designed the system to evaluate execution outcomes against the original objective.&lt;/p&gt;

&lt;p&gt;Orbit — The Workspace&lt;/p&gt;

&lt;p&gt;Orbit is the environment where people and AI agents can work around shared tasks and workflows.&lt;/p&gt;

&lt;p&gt;The bigger vision is to move beyond isolated AI conversations toward a workspace where intent becomes observable, actionable work.&lt;/p&gt;

&lt;p&gt;The core foundation of LocalDesk has already been built. But some of the final components, integrations, and validation work still need to be completed.&lt;/p&gt;

&lt;p&gt;The development budget I had available has now been exhausted. To finish those remaining pieces and bring the complete experience to users, we need additional funding.&lt;/p&gt;

&lt;p&gt;Think of it like unveiling one of the world's most ambitious shopping malls.&lt;/p&gt;

&lt;p&gt;The structure is already standing. The architecture exists. But the final infrastructure, finishing work, safety checks, and operational preparations still require resources before the doors can open.&lt;/p&gt;

&lt;p&gt;That's where LocalDesk stands today. We're not starting from zero. We're preparing for the final stage before the unveiling.&lt;/p&gt;

&lt;p&gt;Our plan is to use a dedicated 168-hour final-build window after the founding event to complete the remaining work and prepare LocalDesk for its next chapter.&lt;/p&gt;

&lt;p&gt;The difficult engineering challenges are exactly why I want to discuss this publicly.&lt;/p&gt;

&lt;p&gt;What Would a Real Completion Contract Look Like?&lt;/p&gt;

&lt;p&gt;Here's an example.&lt;/p&gt;

&lt;p&gt;A human gives an agent a mission:&lt;/p&gt;

&lt;p&gt;"Fix the authentication bug and deploy the update."&lt;/p&gt;

&lt;p&gt;What would count as successful completion?&lt;/p&gt;

&lt;p&gt;I'd want evidence such as:&lt;/p&gt;

&lt;p&gt;The relevant code changes were recorded.&lt;/p&gt;

&lt;p&gt;The expected tests passed.&lt;/p&gt;

&lt;p&gt;The deployment system confirmed the target environment.&lt;/p&gt;

&lt;p&gt;A separate authentication smoke test succeeded.&lt;/p&gt;

&lt;p&gt;The system reported any unresolved risks or limitations.&lt;/p&gt;

&lt;p&gt;Not every task has such a clean success condition.&lt;/p&gt;

&lt;p&gt;And an agent checking its own output can repeat the same mistaken assumptions that produced the original error.&lt;/p&gt;

&lt;p&gt;So verification itself needs to be designed, challenged, and measured.&lt;/p&gt;

&lt;p&gt;For higher-risk actions, human approval may still be essential.&lt;/p&gt;

&lt;p&gt;That's the engineering problem I find most interesting.&lt;/p&gt;

&lt;p&gt;I'm Building This in Public: Mehrad vs Internet&lt;/p&gt;

&lt;p&gt;Alongside LocalDesk, I've started an experiment called Mehrad vs Internet.&lt;/p&gt;

&lt;p&gt;The idea is to bring developers, founders, and early adopters into the process of challenging what we've built and helping shape what comes next.&lt;/p&gt;

&lt;p&gt;We want real Missions.&lt;/p&gt;

&lt;p&gt;Real feedback.&lt;/p&gt;

&lt;p&gt;Real examples of where AI execution breaks.&lt;/p&gt;

&lt;p&gt;And a community that helps shape what trustworthy agentic software should become.&lt;/p&gt;

&lt;p&gt;We're opening founding pre-orders to help fund the final stage of LocalDesk, with a 168-hour final-build window following the founding event.&lt;/p&gt;

&lt;p&gt;It's ambitious, and the unfinished engineering challenges are real.&lt;/p&gt;

&lt;p&gt;I'm documenting the journey rather than pretending the hardest problems have already been solved.&lt;/p&gt;

&lt;p&gt;If you're curious, you can explore the product demo, the founding experiment, and the pre-order details here:&lt;/p&gt;

&lt;p&gt;Explore LocalDesk and Mehrad vs Internet&lt;/p&gt;

&lt;p&gt;Now I Want to Hear From Developers&lt;/p&gt;

&lt;p&gt;Let's make this a real technical discussion.&lt;/p&gt;

&lt;p&gt;Imagine an AI agent has access to your repository, cloud infrastructure, database, and external APIs.&lt;/p&gt;

&lt;p&gt;It performs a complex task and tells you:&lt;/p&gt;

&lt;p&gt;"Done."&lt;/p&gt;

&lt;p&gt;What evidence would you need before believing it?&lt;/p&gt;

&lt;p&gt;Would you trust another LLM acting as a verifier?&lt;/p&gt;

&lt;p&gt;Would you require deterministic tests?&lt;/p&gt;

&lt;p&gt;Would you insist on independent observation of the external system?&lt;/p&gt;

&lt;p&gt;And for which actions would you never remove human approval?&lt;/p&gt;

&lt;p&gt;I'm especially interested in hearing from developers who have built production agents and watched them fail in unexpected ways.&lt;/p&gt;

&lt;p&gt;Because I don't think the future belongs to the agents that sound the smartest.&lt;/p&gt;

&lt;p&gt;I think it belongs to the systems that can prove their work.&lt;/p&gt;

&lt;p&gt;Let's debate it.&lt;/p&gt;

&lt;p&gt;Mehrad Arbab&lt;br&gt;
Founder &amp;amp; CEO, LocalDesk&lt;br&gt;
Building in public: #MehradVsInternet&lt;/p&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>discuss</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
