I'm 19. I've built the foundation of an AI-native workspace called LocalDesk.
And the deeper I go into building AI agents, the more convinced I become that we're asking the wrong question.
Everyone wants more autonomous agents.
Agents that code. Agents that browse. Agents that execute workflows. Agents that operate entire businesses.
But there's a much more important question:
When an AI agent tells you it's done, how do you know it's telling the truth?
Not whether the model is lying intentionally.
Whether the work actually happened.
The "Done" Problem
Imagine asking an AI agent to deploy an application.
It modifies your code, pushes a commit, calls the deployment API, and responds:
"Successfully deployed."
Sounds great.
Except the deployment failed.
Or the wrong environment was updated.
Or the application started successfully but the authentication flow broke.
The agent completed a sequence of actions.
But it didn't necessarily complete your objective.
That distinction matters more as agents receive greater permissions.
More Intelligence Isn't the Whole Solution
We have increasingly capable language models.
They can reason through problems, generate software, and coordinate tools.
But reliability in the real world requires more than generating the next intelligent action.
I think we need a clearer separation between four concepts:
Intent — What does the human actually want?
Execution — What actions did the agent perform?
Verification — What independent evidence shows the objective was achieved?
Recovery — What happens when reality doesn't match the plan?
Without these distinctions, an agent can look incredibly capable while quietly creating operational risk.
What We've Built at LocalDesk
This is one of the architectural problems that shaped LocalDesk.
We've built the core architecture of an AI-native workspace around three connected systems.
Baymax — Human Intent
Baymax is our conversational AI interface.
The goal is to let humans describe what they want to accomplish without manually coordinating every application and workflow.
Missions — Execution and Verification
Missions is our orchestration layer.
Our architecture separates routing, context, planning, execution, verification, and repair.
Instead of treating a model's response as the final result, we've designed the system to evaluate execution outcomes against the original objective.
Orbit — The Workspace
Orbit is the environment where people and AI agents can work around shared tasks and workflows.
The bigger vision is to move beyond isolated AI conversations toward a workspace where intent becomes observable, actionable work.
The core foundation of LocalDesk has already been built. But some of the final components, integrations, and validation work still need to be completed.
The development budget I had available has now been exhausted. To finish those remaining pieces and bring the complete experience to users, we need additional funding.
Think of it like unveiling one of the world's most ambitious shopping malls.
The structure is already standing. The architecture exists. But the final infrastructure, finishing work, safety checks, and operational preparations still require resources before the doors can open.
That's where LocalDesk stands today. We're not starting from zero. We're preparing for the final stage before the unveiling.
Our plan is to use a dedicated 168-hour final-build window after the founding event to complete the remaining work and prepare LocalDesk for its next chapter.
The difficult engineering challenges are exactly why I want to discuss this publicly.
What Would a Real Completion Contract Look Like?
Here's an example.
A human gives an agent a mission:
"Fix the authentication bug and deploy the update."
What would count as successful completion?
I'd want evidence such as:
The relevant code changes were recorded.
The expected tests passed.
The deployment system confirmed the target environment.
A separate authentication smoke test succeeded.
The system reported any unresolved risks or limitations.
Not every task has such a clean success condition.
And an agent checking its own output can repeat the same mistaken assumptions that produced the original error.
So verification itself needs to be designed, challenged, and measured.
For higher-risk actions, human approval may still be essential.
That's the engineering problem I find most interesting.
I'm Building This in Public: Mehrad vs Internet
Alongside LocalDesk, I've started an experiment called Mehrad vs Internet.
The idea is to bring developers, founders, and early adopters into the process of challenging what we've built and helping shape what comes next.
We want real Missions.
Real feedback.
Real examples of where AI execution breaks.
And a community that helps shape what trustworthy agentic software should become.
We're opening founding pre-orders to help fund the final stage of LocalDesk, with a 168-hour final-build window following the founding event.
It's ambitious, and the unfinished engineering challenges are real.
I'm documenting the journey rather than pretending the hardest problems have already been solved.
If you're curious, you can explore the product demo, the founding experiment, and the pre-order details here:
Explore LocalDesk and Mehrad vs Internet
Now I Want to Hear From Developers
Let's make this a real technical discussion.
Imagine an AI agent has access to your repository, cloud infrastructure, database, and external APIs.
It performs a complex task and tells you:
"Done."
What evidence would you need before believing it?
Would you trust another LLM acting as a verifier?
Would you require deterministic tests?
Would you insist on independent observation of the external system?
And for which actions would you never remove human approval?
I'm especially interested in hearing from developers who have built production agents and watched them fail in unexpected ways.
Because I don't think the future belongs to the agents that sound the smartest.
I think it belongs to the systems that can prove their work.
Let's debate it.
Mehrad Arbab
Founder & CEO, LocalDesk
Building in public: #MehradVsInternet
Top comments (0)