A customer asks where a shipment is, whether it will arrive late, and what should happen if it does. A language model alone can produce a polished reply, but it cannot know the current location, inspect live conditions, or change a booking. An AI agent can do those things because it works in a loop: it chooses an action, asks a runtime to execute it, observes the result, and decides what comes next. Tools provide the actions, memory carries useful context, the Model Context Protocol standardises connections to external systems, and orchestration keeps the whole process moving toward a clear stopping point.
The agent loop turns an answer into a process
A single model call takes an input and produces an output. Even if that output is accurate and detailed, the model has no opportunity to act, inspect what happened, and revise its decision.
An agent adds that opportunity. Given a goal, it decides whether it has enough information to answer or needs an action first. After an action returns a result, the agent treats that result as new context. It can then call another tool, change direction, or finish.
Consider an agent handling a shipment question. It may first request the tracking record. If the record reports a disruption, the agent may inspect a schedule and current operating conditions before forming an answer. Each observation affects the next decision.
An agent repeats decisions and actions until it has enough evidence to answer.
Think of it as a doctor ordering a laboratory test. The doctor decides which test is needed, but the laboratory performs it. The result returns to the doctor, who interprets it and decides whether more investigation is necessary. The useful capability comes from the entire loop, not from either participant acting alone.
This distinction gives engineers a practical test: can an external result come back and change what the system does? If not, the application is using a model, but it is not operating as an agent.
Tools let the agent act without making the model the executor
A tool is an external function, API, or data source made available to the agent. It might retrieve a record, search a document store, calculate a value, update a ticket, or submit a booking.
The model does not execute the tool. It identifies the appropriate tool and produces the arguments for a requested call. The surrounding runtime validates and executes that request, then places the result into the model’s context. The model interprets the result and chooses the next step.
That division of responsibility matters. Saying that “the model queried the database” hides the component that actually controls credentials, network access, input validation, timeouts, and returned data. Those responsibilities belong to the runtime and its integrations.
For example, suppose an agent needs the stock level for a part. The model may request get_stock with a part identifier. The runtime calls the inventory system and returns the current quantity. Only then does the model construct its answer. Without the tool, it could generate a plausible quantity; with the tool, it can reason from the retrieved one.
The model requests an action; the runtime executes it and returns the result.
Tools are especially valuable when facts can change after a model was trained. Account state, inventory, schedules, policies, and reservations all require access to current systems. Tools also turn advice into action: instead of describing how to make a booking, an agent can request the booking and inspect its confirmation.
Access should still be bounded. A tool-enabled agent can only request actions exposed by its runtime, and the runtime remains responsible for deciding what is permitted. The agent loop provides flexibility; it does not remove the need for controlled execution.
MCP standardises the connection to external systems
Connecting one application to one tool can be straightforward. Connecting many AI applications to many databases, file stores, internal APIs, and other systems creates a larger integration problem. If every application needs a custom connector for every system, the work grows with every new pair.
The Model Context Protocol, or MCP, is an open protocol that standardises how an AI application connects models to external tools and data sources. Instead of teaching every application the private interface of every external system, each side implements a shared interface.
For an illustrative example, three applications connected independently to four systems require twelve bespoke integrations. With a common protocol, the three applications and four systems each implement their side of the interface, producing seven implementations. Adding another system then requires one new implementation rather than one connector for every application.
A shared protocol replaces many application-to-system connectors with one interface.
Think of MCP as a common electrical socket. Appliance makers build for the socket rather than negotiating a different connection with every building. The socket does not generate electricity or perform the appliance’s job; it only defines how the connection works.
Likewise, MCP does not replace tools. The database still answers queries, and the booking API still creates bookings. MCP standardises how those capabilities are presented and used. Tools can also work without MCP, particularly when a system has only a small number of stable integrations.
The endpoints provide the simplest boundary test: MCP connects an agent to an external system. It does not define how several agents communicate with one another. Agent-to-agent exchange requires communication patterns such as direct messages, shared state, broadcasts, or queues.
MCP is also neither a model nor a hosted service. It is a protocol that compatible applications and systems can speak. Treating it as a product obscures the integration problem it actually solves.
Memory carries selected context, not the entire past
A model is stateless between calls. If an agent needs earlier information, the application must supply that information again. Memory management is the deliberate choice of what to carry forward, how long to keep it, and what to discard.
Working memory contains material needed for the current task: recent turns, the active goal, intermediate decisions, and the latest tool result. Once the task ends, that state usually loses its value.
Persistent memory holds facts that must survive across separate sessions. Examples include a customer’s communication preference or a durable outcome from an earlier case. This information is stored outside the model and retrieved when it becomes relevant.
Task-local context stays in working memory; durable facts enter persistent memory.
Think of working memory as a desk and persistent memory as a filing cabinet. The desk holds the papers required for the current job. The cabinet holds durable records, but you retrieve only the folder relevant to the work in front of you.
Keeping everything is not a safe default. Re-supplying an entire conversation makes every call carry material that may no longer matter. Useful evidence can become harder to find among old details, while unnecessary retention increases cost and risk.
A strong memory policy therefore includes forgetting. A tool result needed for the next decision belongs in working memory. A lasting preference may justify persistent storage. A one-use verification code should be discarded. A long conversation is often better represented by a selected summary than by a verbatim transcript.
The design question is not “How can the agent remember more?” It is “What must survive, and what must not?”
Orchestration coordinates tools, memory, and specialist agents
Workflow orchestration owns the progress of an agentic task. It sequences steps, routes work, handles failures and retries, and decides when the task is finished. Without that control, an agent may have useful tools and memory yet still repeat actions, stall after an error, or stop before reaching the goal.
A deterministic workflow has developer-defined steps and branches. It is predictable, testable, and auditable, making it suitable when the permitted path must remain stable. A model-driven workflow lets the model choose each next step. It can adapt to unexpected situations, but its path may vary between runs.
For example, a routine record lookup can follow a deterministic sequence: validate the identifier, call the record tool, format the response, and stop. An investigation involving incomplete evidence may benefit from model-driven choices about which source to inspect next. The appropriate design depends on whether variation is useful or a defect.
Some tasks also justify multiple agents. The key test is whether the task spans genuinely distinct specialisms, not whether it merely contains many steps.
Bounded work stays with one agent; distinct specialisms justify coordinated agents.
A supervisor pattern uses a lead agent to divide work among specialists and assemble their findings. A sequential pattern passes each agent’s output to the next stage. A parallel pattern lets independent specialists work separately before their results are merged. A hierarchical pattern adds layers of supervision for deeper decomposition.
These structures do not determine how information travels. A supervisor can contact specialists directly or coordinate through shared state. Agents may broadcast an update to several peers or use a message queue when producers and consumers must remain decoupled. The system pattern is the org chart; the communication pattern is how the memos move.
Multiple agents add model calls, latency, failure modes, and debugging surface. A shipment lookup that only needs a tracking tool should remain with one agent. A disruption analysis requiring separate pricing, customs, and capacity expertise may justify a supervisor and specialists. Good orchestration can route those two cases down different paths instead of forcing every request through the larger system.
Key takeaways
- An agent is defined by its loop: decide, act, observe, and decide again.
- The model requests tool calls; the surrounding runtime executes them.
- Tools give agents access to current information and real actions.
- MCP standardises agent-to-system connections but does not replace tools.
- MCP does not govern communication between agents.
- Working memory serves the current task; persistent memory stores selected durable facts.
- Orchestration controls sequencing, routing, failure handling, and completion.
- One agent with good tools is the default; multiple agents need distinct specialisms to justify their cost.
These concepts explain how an agentic application gets work done, but they do not by themselves choose the right permissions, retention policy, tools, or workflow for a particular system. Those decisions depend on the task’s risks, required evidence, tolerance for variable behavior, and the consequences of an incorrect action.





Top comments (1)
i've noticed that a lightweight tool registry keeps the observation loop under 100 ms, how do you handle memory syncing when agents are sharded across multiple nodes?