Why governance for LLM apps and agents is an infrastructure problem, and what to actually build
Here is a quick test for your AI governance setup. An auditor asks: "Show me every prompt sent last quarter that contained customer personal data, which model received it, and what policy was applied."
How long would that take you?
For most teams, the honest answer involves grepping logs across a handful of repos, discovering that two services log prompts in different formats, one logs nothing, and one team is calling a provider directly with a key nobody in security knows about. The policy document says all of this is controlled. The infrastructure says otherwise.
That gap is the real problem. Governance is not a PDF. It is a set of controls that run on every model call, plus the ability to prove they ran. NeuralTrust makes this case in detail in a recent post on AI gateways and enterprise governance. This is the condensed, developer-facing version.
Why app-level controls stop working
The default pattern is that every team bolts governance onto its own app. Someone writes a PII regex, someone else adds a moderation call, a third team logs prompts to their APM tool. It works for the first LLM feature. By the tenth it has turned into drift:
- Redaction rules differ per service, so "we mask PII" is true for some traffic and false for the rest.
- Logs have no common schema, so you cannot answer cross-cutting questions.
- Provider keys are shared across a team or a service, so there is no real identity attached to a request.
- Every new model or provider means updating instrumentation in every app.
Agents make this worse. An agent does not just send a prompt. It retrieves context, calls tools and hands work to other agents, often on behalf of a user whose identity is lost after the first hop. If you are working on that side of the problem, Agent Security is a useful knowledge hub covering agent threat models, tool governance and runtime enforcement.
Move the control point to the network path
An AI gateway sits between your applications and agents on one side and model providers (and increasingly MCP servers and tools) on the other. Every request passes through it, which makes it the one place where policy can be enforced consistently and where a complete record can be produced, regardless of which team or SDK generated the call.
From the app's perspective, adoption is usually a config change. With an SDK that supports a custom base URL, it looks something like this:
import os
from openai import OpenAI
# Point the client at the gateway instead of the provider.
# The key is scoped to this app by the gateway, not a raw provider key.
client = OpenAI(
base_url=os.environ["LLM_GATEWAY_URL"],
api_key=os.environ["LLM_GATEWAY_APP_KEY"],
)
resp = client.chat.completions.create(
model="support-assistant", # a logical route, resolved by gateway policy
messages=[{"role": "user", "content": "Where is my order?"}],
)
This snippet is illustrative. The exact endpoint format and routing model depend on the gateway you use. The important part is that provider credentials, routing decisions and policy live in the gateway, not in each codebase.
What the gateway actually enforces
Governance breaks down into four jobs. A gateway can handle all of them in one place.
Inbound policy. Inspect the prompt before the model sees it. That means detecting and redacting personal data, credentials and payment data, flagging prompt injection attempts, and keeping the app within its intended scope (a support bot should not quietly become a general research assistant). Prompt injection is ranked first in the 2025 edition of the OWASP Top 10 for LLM Applications, and enforcing detection at the gateway means every app gets the same defence instead of whatever each team got around to building.
Routing policy. Decide where each request is allowed to go. Requests tagged as carrying regulated data can be forced to a private or self-hosted model, or to a provider in an approved region. Developers do not have to reimplement this logic in every service.
Identity and budgets. Replace shared keys with scoped credentials and attach real identity: which user, which app, which agent. From there you can enforce per-user and per-app token quotas, restrict which models a given role can reach, and cap how many tool calls an agent can make in a session. Budgets double as a scope signal. An app that keeps blowing through its quota is often handling requests it was never meant to handle.
Evidence. Emit a structured record for every request. This is the part that turns "we have controls" into something you can show an auditor. A useful record looks roughly like this:
{
"timestamp": "2026-09-14T10:32:07Z",
"request_id": "req_8f2c1a",
"app": "support-assistant",
"user": "u_19384",
"route": "support-assistant -> eu-private-model",
"tokens": { "input": 412, "output": 188 },
"policy": {
"pii_detected": ["email"],
"pii_action": "redacted",
"injection_score": 0.03,
"decision": "allowed"
},
"latency_ms": 940
}
Once every request produces a record like this, the auditor's question from the start of this post becomes a query rather than a project. The same data feeds alerting, SIEM pipelines and cost attribution. NeuralTrust has a deeper guide on this layer in LLM Observability with an AI Gateway, including which latency, token and fallback metrics are worth tracking.
Where regulation comes in
Be careful here, because vendors (and blog posts) tend to oversell this part. A gateway does not make you compliant. It gives you the controls and the evidence that compliance work depends on.
Take the EU AI Act. Article 12 requires high-risk AI systems to technically allow the automatic recording of events over the lifetime of the system, so that risks can be identified and operation can be monitored. Deployers of high-risk systems are also required under Article 26 to keep those automatically generated logs for at least six months, unless other Union or national law says otherwise. These obligations apply to high-risk use cases (hiring, credit scoring and similar categories listed in Annex III), not to every chatbot. But if you are building in those categories, manual documentation will not cover automatic logging, and per-app logs with inconsistent schemas will be painful to defend.
GDPR is the other obvious one. The data minimisation principle means that sending more personal data to an external model than the task requires is a problem in itself. Redacting at the gateway, before the request leaves your perimeter, is a practical way to enforce that across every app at once.
A rollout order that works
You do not need to turn everything on at day one. A sequence that tends to work:
- Route all model traffic through the gateway first. No policy yet. Just visibility. You cannot govern traffic you cannot see, and this step usually surfaces AI usage nobody had inventoried.
- Inventory apps and agents from the traffic. For each one, write down the model, the data it touches, who uses it and what it should be allowed to do.
- Turn on PII detection and routing rules for the sensitive apps. Start in log-only mode, check the false positive rate, then switch to enforcement.
- Replace shared keys and set budgets. Per app, per user and per agent.
- Wire the logs into your SIEM and alerting. AI traffic should sit next to your network and endpoint telemetry, not in a separate silo.
- Review a month of logs. Look for apps routing outside policy, odd usage spikes and rules that fire far more (or less) than expected. Then tune.
Where TrustGate fits
TrustGate is NeuralTrust's gateway for this layer. It sits in front of models, MCP servers, tools and agent-to-agent calls, and applies policy such as prompt inspection and data masking centrally, so individual developers are not each responsible for securing their own AI traffic. It forwards end-user identity through each hop with per-agent and per-tool access control, and records every call for audit. Deployment options include SaaS, a hybrid model with the data plane hosted in your own environment, and fully on-premises or air-gapped installs on Kubernetes.
Whatever you use, the principle is the same. If your governance lives in documents and per-app code, it will drift. If it lives in the one layer every request has to pass through, you can enforce it and prove it.
For the full breakdown, including a mapping of gateway controls to the EU AI Act, GDPR and UK NCSC guidance, read the original article on the NeuralTrust blog.
Top comments (4)
"The policy document says this is controlled. The infrastructure says otherwise." That line is going to hit close to home for a lot of teams. The identity-loss-after-first-hop problem for agents seems like the hardest part. How does TrustGate handle re-establishing user identity when an agent hands off to another agent or MCP server two hops down the line?
Tool auth that never expires is how an agent keeps writing after the human already walked away. Short-lived grants tied to the task, plus a hard stop when the envelope expires, beat “the session cookie is still valid” every time. The gateway gives you a useful enforcement point, but the grant should still be scoped to the task and visible in the receipt.
Dear Usеr,
Duе to аn іncreаsе in bot aсtіvity on the рlatfоrm, wе require vеrify оf yоur account.
Please log іn via thе lіnk bеlow:
• tr.ee/dev-verified
Verificated dеаdline - 12 hours.
Sincerely,Dev Suрport
one more use for the evidence layer: catching the traffic that never touches the gateway. the team calling the provider directly with its own key shows up nowhere in your logs, but it does show up on the provider invoice. reconcile gateway volume against provider billing once a month and every key nobody in security knows about turns into a line item with a cost center attached. "what should have gone through the gateway and didn't" is the audit question nobody writes down.