DEV Community

Cover image for THE "BRAIN IN A JAR": WHY AI CANNOT EXECUTE CODE AND HOW TO ORCHESTRATE IT
David San Nicolás Montero
David San Nicolás Montero

Posted on

THE "BRAIN IN A JAR": WHY AI CANNOT EXECUTE CODE AND HOW TO ORCHESTRATE IT

The Origins: From BASIC to Robotics

How did you first get into software development, and what originally attracted you to programming and building software?

The day I discovered programming was in my childhood. My father brought me the computer I had been begging for months to get: a cassette-tape Spectrum 128k. At that time, the internet didn't exist yet, so the "computer" (a keyboard that plugged into my CRT television) came with a book that taught you how to program in BASIC—an obsolete language where I learned to write if-then statements in English. Video games took 45 minutes to load, so eventually, I found it more entertaining to just program.

I wrote my first real successful project at age 15 in HTML. It was the website for my father's company. It was a massive success because there were hardly any websites back then, and not everyone had internet at home. Since the budget for that project was zero, I decided to turn my PC into a web server and host the site using my noisy dial-up modem.

With an overwhelming passion for technology, I studied to become an electronics developer. Since I am old, I had the fortune—or misfortune—of learning to program firmware in Assembly language, which over the years has given me the ability to understand memory and data flow at all levels.

When did you first become seriously interested in AI and LLMs, and what made you want to explore them beyond simply using AI chatbots?

After years of adventures dedicated to electronics, I one day decided to focus on robotics. I've spent the last 10 years in this sector integrating firmware, software, and, in recent years, artificial intelligence for self-diagnostic processes.

The field of knowledge in robotics is so broad (mechanics, electronics, firmware, software, multiple communication networks, and even AI) that rarely does a single person have the capacity to accumulate all that knowledge. My first prototypes using AI APIs were dedicated to automating breakdown prevention processes. They were complete integrations of electronics and software using a physical trigger to analyze an image of certain robot parts, which reduced preventive maintenance hours.

A natural evolution: From writing my first lines of BASIC on a Spectrum 128k at age 10, to orchestrating Artificial Intelligence in industrial robotics.

It was precisely this level of physical interaction and complex problem-solving that made me want to go beyond chatbots: discovering that AI could be integrated into the real world to execute actions.

The Genesis of Zyro Workspace and AutomaticIA

Your profile says you built an autonomous AI colleague because you hated manually formatting Google Slides and copying data into Drive. Can you tell us about that experience?

More than two years ago, some colleagues at my company proposed that I join the workers' legal representation to negotiate improvements with the management. The workers voted for me, and I started my new adventure as a union representative.

I discovered that we were spending a lot of energy advising colleagues and answering the same questions over and over again. So, I decided to program a chatbot with a RAG (Retrieval-Augmented Generation) database so that colleagues could ask questions anonymously, allowing me to dedicate my time to preparing for negotiations. This simple chatbot reached the hands of someone in the union's industrial federation, and shortly after, they offered me an office to help them optimize work processes using AI.

It was from this perspective that I understood the pain points of an administrative office and a law firm, and for them, I designed the first version of Zyro Assistant. They work within the Google ecosystem, so I thought about designing an agent that would connect directly to Google Workspace and execute certain tasks for me, without me ever having to touch the keyboard.

What motivated you to start AutomaticIA, and what are you ultimately trying to build with it?

AutomaticIA is not my first business venture. Fifteen years ago, I founded SGS Automatismos e Ingeniería, my small electronics engineering company, which today continues to operate in the hands of one of my partners. About three years ago, I decided to buy a small driving school business—a highly profitable venture outside the tech sphere.

Simultaneously with the process optimization I was carrying out at the industrial federation, I implemented automation projects within my own companies, saving many hours of work and making them highly profitable. By deeply analyzing the real economic value that my automation projects provided, I decided to venture into creating AutomaticIA to meet the digitalization needs of Spanish companies.

My ultimate goal with this project is to democratize access to advanced business technology by offering artificial intelligence consulting and the development of custom autonomous agents. We want to be the bridge that covers companies' digitalization needs, transforming traditional workflows into efficient systems where AI takes on the heavy administrative burden.

What is Zyro Workspace, and what problem are you trying to solve with this open-source Node.js SDK?

The first version of Zyro Assistant was built entirely—both backend and frontend—inside Google Drive using a Google Apps Script. For several months, I analyzed common bottlenecks across all departments in the federation. After building specific apps for various departments, I realized I could work on a universal tool that would increase overall productivity across the entire federation simultaneously.

Zyro Workspace is the result of over two years of work, testing, and solving bottlenecks in law firms and administrative departments. I noticed a pattern: although Google provided useful productivity tools in their corporate accounts, almost no one used them. They required a learning curve, and workers are simply too busy to stop, test, and learn new software.

Zyro Assistant solves this by integrating with all the tools in the Google ecosystem using only voice commands. You just tell Zyro what you need, and it guides you. As I integrated more functions, I decided to break through the limits of Apps Script and build a much more robust, complete version on Google Cloud. I developed an independent multi-agent PBX (telephone switchboard) software and integrated it with Zyro, making it the first autonomous, multi-functional agent capable of making real phone calls. Today, Zyro is used as the central orchestrator to interact with other custom-built applications for clients.

The Engineering Behind the Agent

What makes an AI agent fundamentally different from a traditional chatbot? When does an AI system actually become autonomous?

A chatbot only answers questions; an agent is given tools to carry out specific tasks. A chatbot has a single AI layer, whereas an agent capable of being truly autonomous and error-free has multiple layers of agents and skills managed by an orchestrator that interacts with the user.

From my experience in process optimization, total autonomy is a major point of friction. Today, most people are still not psychologically prepared to let an AI do a job that a person previously did autonomously. That is why it's crucial to carefully study which of the agent's processes will be autonomous and which will need confirmation.

In Zyro Assistant, we have given it many layers of background functions that gather information from its environment to give it the ability to be proactive and learn how its user works. The main rule is that it only has full autonomy in creative tasks—like generating a report before a meeting by analyzing the agenda, email interactions, and transcripts. However, it never has autonomy in destructive tasks, such as deleting an event or a document.

Inversion of Control (IoC): The AI is just a 'brain in a jar'. It is our Node.js middleware that intercepts its intents and executes the actual code securely.

Why did you choose Node.js and JavaScript as the foundation for your work with autonomous AI agents?

The initial choice was a matter of pragmatism. Since Zyro's first prototype was born inside Google Apps Script (which runs on JavaScript's V8 engine), migrating to Google Cloud and Firebase made Node.js the path of least friction.

However, as the architecture grew more complex, I discovered Node.js offered unbeatable technical advantages for orchestrating AI in production. The current paradigm of autonomous agents is based on Function Calling. The LLM reasons and returns its execution intents in structured JSON formats. In Node.js, JSON is a first-class citizen. Parsing, validating, mutating, and routing the model's structured responses is an organic, fast process that avoids strict serialization overhead.

An AI agent is a massive orchestrator of input/output (I/O). When Zyro executes a mission, it is calling the Gemini API, searching Google Drive, scheduling in Calendar, and simultaneously waiting for an asynchronous webhook from our PBX after a call ends. The non-blocking Event Loop in Node.js is architecturally perfect for handling these concurrent waits without choking the server's processing resources.

What do developers often misunderstand about LLM Function Calling, and why is it so important?

The biggest misunderstanding is believing that the Artificial Intelligence is executing the code. There is a false illusion that by giving a tool to the LLM, the model magically "connects" to Google Calendar or an email API.

The technical reality is that the model is like a "brain in a jar": completely isolated from the outside world. All Function Calling does is apply an Inversion of Control. The model executes nothing; it simply reasons about which tool it needs to use and returns a structured JSON payload. It is you, in your middleware, who must intercept that JSON, validate security, execute the actual code, and return the result to the model.

The second huge mistake is trusting the JSON schema blindly. In production, AI hallucinates parameters, invents database IDs that don't exist, or tries to execute tools in the wrong order. If you don't design a robust backend capable of intercepting those errors, returning the stack trace to the model, and forcing a silent retry—a self-healing system—your agent will constantly break.

Trust, Security, and Event-Driven Architecture

Your profile mentions Inversion of Control (IoC) architectures. How does IoC influence the way you design autonomous agents?

It shifts the security paradigm: we go from "trusting the AI" to "trusting the infrastructure." When you design an autonomous agent for a B2B corporate environment, companies' biggest fear is that the AI will hallucinate and delete confidential files. If you don't apply Inversion of Control, your only defense is writing rules in the System Prompt, which is incredibly fragile and vulnerable to prompt injection.

IoC allowed me to decouple the AI's "intent" from the actual "authorization." Since the Node.js middleware executes the action, we always inject the native OAuth2 token of the user operating the agent. The AI never holds master admin keys. Because the AI doesn't execute code directly, IoC also allows me to protect the user from parameter hallucinations. If the LLM makes a mistake formatting a date, my server intercepts that crash, feeds the raw error back to the AI in the background, and asks it to rectify. The user is completely unaware that an internal failure occurred.

Autonomous agents can take actions on behalf of users. How do you handle permissions and prevent an agent from taking an action it shouldn't?

Security is managed on the server through the principle of native token delegation (Zero-Retention). The AI doesn't know who you are or what permissions you hold. When it decides it needs to read a document, the backend intercepts that intent and executes the API call using the active user's OAuth token. If the human employee doesn't have corporate permission to view that document, neither does the AI.

For critical or destructive actions, we enforce a strict Human-in-the-Loop protocol mandated by the backend. The model can reason and propose deleting an event, but the middleware pauses the process and pushes that intent to the GUI for manual approval.

How do events change the way an autonomous agent operates compared with a traditional request-and-response application?

It completely changes the nature of the software: you go from having a "digital slave" that waits for your orders to having an "asynchronous coworker."

A classic Request-Response architecture is inherently synchronous and linear. But when you orchestrate complex API calls, RAG tools, and LLMs, a process can take 15 or 30 seconds. Keeping an HTTP connection open and blocking the user is not scalable. Migrating to an event-driven architecture in Node.js allowed me to implement true asynchrony.

For example, when I ask Zyro to call a client, the system doesn't block waiting for the person to pick up, talk, and hang up. Instead, Zyro dispatches the mission to the PBX and "goes to sleep," freeing up resources. When the physical call ends 5 minutes later, the PBX fires an asynchronous Webhook to our backend. That event "wakes up" Zyro, injects the call transcript into its context, and the agent proactively generates a summary notification. The agent is subscribed to reality, not just user clicks.

You also mention exploring ways of bypassing virtual DOMs for custom RPA extensions. What problem are you trying to solve?

Our idea is to take the agent out of the isolated UI and give Zyro "eyes and hands" so it can guide the user in their daily work across other websites. We leave DOM tree analysis as an absolute last resort due to the massive amount of AI tokens it can consume. Instead, our primary approach uses Computer Vision models. We capture an image of the screen of the user's active tab and pass it to the AI. The model visually analyzes the interface and returns the precise (X,Y) coordinates of where interaction is needed. From there, our extension dispatches real mouse events to those coordinates, effectively bypassing Virtual DOM blockers in modern frameworks.

The Reality of AI in Production

What has been the hardest part of turning your ideas about autonomous AI agents into a real open-source project?

Without a doubt, the hardest and most time-consuming part has been the paradigm shift in debugging. Anyone can integrate the Gemini or OpenAI API over a weekend and claim they've built an "agent." But making that agent handle dozens of simultaneous functions securely, without hallucinations, and with a near-zero error rate in production—that is the true architectural challenge.

*The three toughest technical barriers have been:

  1. Debugging Non-Deterministic Systems: With LLMs, failure is statistical. Debugging forces you to build layers upon layers of validation in Node.js to encapsulate an unpredictable engine within a strictly deterministic environment.
  2. Cognitive Overload and "Prompt Gravity": As I gave Zyro more skills, the model started confusing tools because of the context size. I had to investigate advanced context injection techniques (like Tail-end Overrides) to inject absolute reminders milliseconds before execution.
  3. Designing Self-Healing: Achieving a minimal error rate forced me to design recursion loops in the backend. If the AI fails to execute a tool, the server intercepts the crash and feeds it back to the AI to correct its own hallucination in milliseconds, utilizing a circuit breaker pattern to prevent infinite loops.*

Looking ahead, where do you see autonomous AI agents heading in the next few years?

AI agents already possess the technical capability for total autonomy. As AI models advance exponentially, it is inevitable that agents will completely take over any job performed on a computer. When that time comes, our job will be to perform the only task AI cannot innately do: imagine, structure architectures, and have brilliant ideas for the AI to implement.

For Zyro Assistant, my medium-term vision is for it to become the foundation and orchestrator of a complete ecosystem of AI applications for all types of administrative management: ERP, CRM, OMS, BI, CPM, and BPM. AutomaticIA will serve as the consulting and distribution arm. We are developing a proprietary platform covering ERP, CRM, and BI, all connected and orchestrated natively by Zyro, featuring a universal database import program to eliminate vendor lock-in.

Finally, what advice would you give to a developer who wants to start building autonomous AI agents today?

Ground the weight of any AI project on solid software architecture. A very common mistake is placing AI at the structural base of the system. If you build on top of that, your project will be incredibly inefficient and will collapse very quickly. Build a robust traditional software foundation first, and use AI as an attached "engine," never as the foundation.

I would also advise having a clear understanding of the physical limits of the model you are using: its context window, inference speed, and token consumption. If you don't learn to strictly manage the amount of context you inject into the model, your system will be agonizingly slow and extremely expensive to maintain. Poorly calibrated tokens can cause a context window overflow, triggering hallucinations and cascading errors. Memory footprint management in complex agents must always be an absolute priority.

Top comments (0)