I have spent about ten years in Microsoft's ecosystem, and most recently a lot of that time has gone into building and governing agents in Copilot Studio. In that world, an agent is something you configure. You describe its instructions, connect knowledge and tools, pick a channel, and the platform supplies everything underneath.
Then a simple question started bothering me: what is actually underneath? If I could not answer that, I could not honestly assess an agent built somewhere else. And in most enterprises I work with, agents are being built somewhere else. A data team uses an open-source framework, a business unit pilots on Google Cloud, and the central platform team is still expected to produce one inventory and one risk picture.
So I started building agents in code with Google's Agent Development Kit (ADK) and Gemini. This post is not the step-by-step setup; that is a separate, longer walkthrough. This is the mental model I wish I had on day one, written for developers and for anyone who has to govern what developers build.
Two ways to get an agent
There are broadly two ways an organization ends up with an AI agent.
The first is to configure one on a managed platform. Microsoft describes Copilot Studio as a graphical, low-code studio for building and managing agents and workflows, with analytics, evaluations, and an administration layer for inventory, role-based access, and cost management (Copilot Studio overview). You make decisions about behavior. The platform makes most decisions about runtime, hosting, and identity integration.
The second is to engineer one with a framework. Google describes ADK as an open-source, code-first Python framework that is model-agnostic and deployment-agnostic (google-adk on PyPI). You write the agent as software. The framework gives you a runtime and conventions, but where it runs, how users sign in, and how it is monitored are now your architecture decisions.
Neither is better in the abstract. The useful way to compare them is to ask where responsibility sits.
That last row is the one I keep coming back to. A low-code agent at least appears in an admin center. A code-first agent can live in a repository and on a laptop indefinitely, with a working API key and no record anywhere that it exists.
Local agent, remote model
The first concept that clicked for me was that "the agent" and "the model" are in different places.
My development environment is a Fedora Linux workstation with a 2 GB graphics card, and my first test request to Gemini 3.1 Flash-Lite came back with a valid response. That card could not host a modern large language model, and it did not need to. The agent process runs on my machine: the Python runtime, the instructions, the tool functions, and the conversation state. The model runs in Google's cloud and is reached over HTTPS through the Gemini API. My workstation needs a network connection and a Python interpreter, not a data-center GPU.
One turn of a conversation looks roughly like this:
You ──message──▶ Agent runtime (local)
│ sends: your message + instructions + tool schemas
▼
Gemini API (remote) ──▶ "call tool X with these arguments"
│
Agent runtime runs tool X locally, sends the result back
▼
Gemini API (remote) ──▶ final answer
│
You ◀──reply─────┘
This split is easy to state and easy to forget. Every instruction, every user message, and every tool result goes to the model provider on every call. If a tool returns customer records, those records leave your environment in the next request.
Which rules govern that data depends on how you pay. Google's Gemini API terms separate unpaid use, where submitted content may be used to improve Google's products and may be read by human reviewers, from paid use billed through a Cloud project, where Google states prompts and responses are not used for product improvement. The terms explicitly say not to send sensitive, confidential, or personal information to the unpaid services. For a learning project, that means synthetic and public data only. For an enterprise, it means data classification has to happen before a tool is written, not after.
What a framework adds, and what it does not
If a single SDK call already reaches Gemini, why use an agent framework at all?
Because a raw model call gives you one request and one response. Everything that makes something an agent sits around that call. The loop that lets the model request a tool, run it, and continue. Generating tool schemas from ordinary functions. Holding conversation state between turns. An event stream you can inspect when behavior surprises you. Evaluation commands, and packaging for deployment. Without a framework, all of that is code you write and maintain. ADK supplies it as conventions: in its quickstart, an agent is a Python module that defines a root_agent with a model, an instruction, and a list of tools (ADK Python quickstart).
The vocabulary is worth separating, because the word "agent" hides several jobs:
• The model reasons. It reads the context and either answers or asks for a tool.
• The instructions are standing orders sent with every request. They shape behavior; they do not enforce it.
• The tools are your functions. The model can only ask for them; the runtime decides whether to run them.
• The runtime orchestrates: it builds each request, executes tools, and keeps session state.
Just as important is what a first agent does not have. It has no long-term memory unless you attach a memory service. It does no planning beyond what the model does within a turn. It has no access to anything except the tools you give it. And it has no enterprise security by default. ADK's own CLI reference states that its local web UI and API server expose unauthenticated endpoints, and tells you to put your own authentication in front before serving other users (ADK CLI reference).
A working response proves that a pipe exists. It does not prove that anything should flow through it yet.
The questions that do not change
Here is the part that surprised me most. Moving from configured agents to engineered agents changed almost every implementation detail, and none of the governance questions.
Whose identity does the agent use? Every agent eventually calls something, and "as the signed-in user" versus "as the agent" decides what data it can reach. During development that is a personal API key. In production it should be a dedicated service identity with least privilege.
What data can it see, and where does that data go? With a hosted model, the answer to the second half is always "to the provider," so the first half needs a classification decision.
Who owns it? An agent without a named business owner and technical owner is a liability with a working credential.
How do you know it still behaves? Model output is not deterministic, and model identifiers change status over time. Pin stable versions and keep a repeatable evaluation set.
Who pays, and who notices? Each tool call is another model round trip that resends the context, so cost scales with agent design, not just user count. Someone has to watch tokens, errors, and latency.
In Copilot Studio, part of the answer to each question exists before you build. With a code-first framework, the answers live in whatever architecture you put around the code. Deploy into a well-governed cloud project with managed identities and central logging, and the controls can be strong. Run it from a laptop with a personal key, and there are effectively none. The framework itself is neutral.
That is why I think enterprises need one governance model that spans platforms rather than one per vendor. I am exploring that idea in a side project I call EnterpriseAgentGuard. It is only an idea and a project folder today, and I will write about it as it takes shape.
Sources
• Microsoft Copilot Studio overview
• google-adk on PyPI
• ADK Python quickstart
• ADK CLI reference
• Gemini API Additional Terms of Service

Top comments (0)