<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Alessandro Pignati</title>
    <description>The latest articles on DEV Community by Alessandro Pignati (@alessandro_pignati).</description>
    <link>https://dev.to/alessandro_pignati</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3663725%2F49945b08-2d78-4735-af16-07e967b19122.JPG</url>
      <title>DEV Community: Alessandro Pignati</title>
      <link>https://dev.to/alessandro_pignati</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/alessandro_pignati"/>
    <language>en</language>
    <item>
      <title>Your AI Policy Doesn't Run in Production. Your Gateway Does.</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Mon, 28 Sep 2026 14:40:07 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/your-ai-policy-doesnt-run-in-production-your-gateway-does-jgj</link>
      <guid>https://dev.to/alessandro_pignati/your-ai-policy-doesnt-run-in-production-your-gateway-does-jgj</guid>
      <description>&lt;h2&gt;
  
  
  Why governance for LLM apps and agents is an infrastructure problem, and what to actually build
&lt;/h2&gt;

&lt;p&gt;Here is a quick test for your AI governance setup. An auditor asks: "Show me every prompt sent last quarter that contained customer personal data, which model received it, and what policy was applied."&lt;/p&gt;

&lt;p&gt;How long would that take you?&lt;/p&gt;

&lt;p&gt;For most teams, the honest answer involves grepping logs across a handful of repos, discovering that two services log prompts in different formats, one logs nothing, and one team is calling a provider directly with a key nobody in security knows about. The policy document says all of this is controlled. The infrastructure says otherwise.&lt;/p&gt;

&lt;p&gt;That gap is the real problem. Governance is not a PDF. It is a set of controls that run on every model call, plus the ability to prove they ran. NeuralTrust makes this case in detail in a recent post on AI gateways and enterprise governance. This is the condensed, developer-facing version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why app-level controls stop working
&lt;/h2&gt;

&lt;p&gt;The default pattern is that every team bolts governance onto its own app. Someone writes a PII regex, someone else adds a moderation call, a third team logs prompts to their APM tool. It works for the first LLM feature. By the tenth it has turned into drift:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Redaction rules differ per service, so "we mask PII" is true for some traffic and false for the rest.&lt;/li&gt;
&lt;li&gt;Logs have no common schema, so you cannot answer cross-cutting questions.&lt;/li&gt;
&lt;li&gt;Provider keys are shared across a team or a service, so there is no real identity attached to a request.&lt;/li&gt;
&lt;li&gt;Every new model or provider means updating instrumentation in every app.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents make this worse. An agent does not just send a prompt. It retrieves context, calls tools and hands work to other agents, often on behalf of a user whose identity is lost after the first hop. If you are working on that side of the problem, &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;Agent Security&lt;/a&gt; is a useful knowledge hub covering agent threat models, tool governance and runtime enforcement.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move the control point to the network path
&lt;/h2&gt;

&lt;p&gt;An AI gateway sits between your applications and agents on one side and model providers (and increasingly MCP servers and tools) on the other. Every request passes through it, which makes it the one place where policy can be enforced consistently and where a complete record can be produced, regardless of which team or SDK generated the call.&lt;/p&gt;

&lt;p&gt;From the app's perspective, adoption is usually a config change. With an SDK that supports a custom base URL, it looks something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="c1"&gt;# Point the client at the gateway instead of the provider.
# The key is scoped to this app by the gateway, not a raw provider key.
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LLM_GATEWAY_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;LLM_GATEWAY_APP_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;support-assistant&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# a logical route, resolved by gateway policy
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Where is my order?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This snippet is illustrative. The exact endpoint format and routing model depend on the gateway you use. The important part is that provider credentials, routing decisions and policy live in the gateway, not in each codebase.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the gateway actually enforces
&lt;/h2&gt;

&lt;p&gt;Governance breaks down into four jobs. A gateway can handle all of them in one place.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inbound policy.&lt;/strong&gt; Inspect the prompt before the model sees it. That means detecting and redacting personal data, credentials and payment data, flagging prompt injection attempts, and keeping the app within its intended scope (a support bot should not quietly become a general research assistant). Prompt injection is ranked first in the 2025 edition of the &lt;a href="https://genai.owasp.org/llmrisk/llm01-prompt-injection/" rel="noopener noreferrer"&gt;OWASP Top 10 for LLM Applications&lt;/a&gt;, and enforcing detection at the gateway means every app gets the same defence instead of whatever each team got around to building.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Routing policy.&lt;/strong&gt; Decide where each request is allowed to go. Requests tagged as carrying regulated data can be forced to a private or self-hosted model, or to a provider in an approved region. Developers do not have to reimplement this logic in every service.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Identity and budgets.&lt;/strong&gt; Replace shared keys with scoped credentials and attach real identity: which user, which app, which agent. From there you can enforce per-user and per-app token quotas, restrict which models a given role can reach, and cap how many tool calls an agent can make in a session. Budgets double as a scope signal. An app that keeps blowing through its quota is often handling requests it was never meant to handle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Evidence.&lt;/strong&gt; Emit a structured record for every request. This is the part that turns "we have controls" into something you can show an auditor. A useful record looks roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-14T10:32:07Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"request_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"req_8f2c1a"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"app"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-assistant"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"u_19384"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"route"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"support-assistant -&amp;gt; eu-private-model"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;412&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;188&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"policy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"pii_detected"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"pii_action"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"redacted"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"injection_score"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.03&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"decision"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"allowed"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"latency_ms"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;940&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once every request produces a record like this, the auditor's question from the start of this post becomes a query rather than a project. The same data feeds alerting, SIEM pipelines and cost attribution. NeuralTrust has a deeper guide on this layer in &lt;a href="https://neuraltrust.ai/blog/ai-gateway-llm-observability" rel="noopener noreferrer"&gt;LLM Observability with an AI Gateway&lt;/a&gt;, including which latency, token and fallback metrics are worth tracking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where regulation comes in
&lt;/h2&gt;

&lt;p&gt;Be careful here, because vendors (and blog posts) tend to oversell this part. A gateway does not make you compliant. It gives you the controls and the evidence that compliance work depends on.&lt;/p&gt;

&lt;p&gt;Take the EU AI Act. &lt;a href="https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-12" rel="noopener noreferrer"&gt;Article 12&lt;/a&gt; requires high-risk AI systems to technically allow the automatic recording of events over the lifetime of the system, so that risks can be identified and operation can be monitored. Deployers of high-risk systems are also required under Article 26 to keep those automatically generated logs for at least six months, unless other Union or national law says otherwise. These obligations apply to high-risk use cases (hiring, credit scoring and similar categories listed in Annex III), not to every chatbot. But if you are building in those categories, manual documentation will not cover automatic logging, and per-app logs with inconsistent schemas will be painful to defend.&lt;/p&gt;

&lt;p&gt;GDPR is the other obvious one. The data minimisation principle means that sending more personal data to an external model than the task requires is a problem in itself. Redacting at the gateway, before the request leaves your perimeter, is a practical way to enforce that across every app at once.&lt;/p&gt;

&lt;h2&gt;
  
  
  A rollout order that works
&lt;/h2&gt;

&lt;p&gt;You do not need to turn everything on at day one. A sequence that tends to work:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Route all model traffic through the gateway first.&lt;/strong&gt; No policy yet. Just visibility. You cannot govern traffic you cannot see, and this step usually surfaces AI usage nobody had inventoried.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inventory apps and agents from the traffic.&lt;/strong&gt; For each one, write down the model, the data it touches, who uses it and what it should be allowed to do.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Turn on PII detection and routing rules for the sensitive apps.&lt;/strong&gt; Start in log-only mode, check the false positive rate, then switch to enforcement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replace shared keys and set budgets.&lt;/strong&gt; Per app, per user and per agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Wire the logs into your SIEM and alerting.&lt;/strong&gt; AI traffic should sit next to your network and endpoint telemetry, not in a separate silo.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Review a month of logs.&lt;/strong&gt; Look for apps routing outside policy, odd usage spikes and rules that fire far more (or less) than expected. Then tune.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Where TrustGate fits
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;TrustGate&lt;/a&gt; is NeuralTrust's gateway for this layer. It sits in front of models, MCP servers, tools and agent-to-agent calls, and applies policy such as prompt inspection and data masking centrally, so individual developers are not each responsible for securing their own AI traffic. It forwards end-user identity through each hop with per-agent and per-tool access control, and records every call for audit. Deployment options include SaaS, a hybrid model with the data plane hosted in your own environment, and fully on-premises or air-gapped installs on Kubernetes.&lt;/p&gt;

&lt;p&gt;Whatever you use, the principle is the same. If your governance lives in documents and per-app code, it will drift. If it lives in the one layer every request has to pass through, you can enforce it and prove it.&lt;/p&gt;

&lt;p&gt;For the full breakdown, including a mapping of gateway controls to the EU AI Act, GDPR and UK NCSC guidance, read the original article on the &lt;a href="https://neuraltrust.ai/blog/ai-gateway-enterprise-governance" rel="noopener noreferrer"&gt;NeuralTrust blog&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>agents</category>
    </item>
    <item>
      <title>Your LLM App Passed Every Security Scan. It Still Leaked Data Through a Calendar Invite.</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Sat, 12 Sep 2026 16:09:13 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/your-llm-app-passed-every-security-scan-it-still-leaked-data-through-a-calendar-invite-4mln</link>
      <guid>https://dev.to/alessandro_pignati/your-llm-app-passed-every-security-scan-it-still-leaked-data-through-a-calendar-invite-4mln</guid>
      <description>&lt;h3&gt;
  
  
  What actually matters when you're picking AI security tooling in the UK market right now
&lt;/h3&gt;

&lt;p&gt;Here's a scenario that's becoming routine. A support agent built on an LLM has access to a ticketing tool and a calendar. Someone emails in a request, the agent reads it, and buried in the email signature is a string of text instructing the model to forward customer records to an external address. No malware, no exploited CVE, no firewall rule that would ever catch it. The agent just did what the text told it to do, because from the model's point of view there's no reliable line between "instructions from my operator" and "text I happened to read."&lt;/p&gt;

&lt;p&gt;That's indirect prompt injection, and it's the reason a WAF or a SIEM won't save you here. Those tools inspect network traffic and system logs. They have no concept of a prompt, a completion, or a tool call, so an attack that lives entirely inside natural language sails straight through.&lt;/p&gt;

&lt;h2&gt;
  
  
  The threat model changed, the tooling didn't
&lt;/h2&gt;

&lt;p&gt;If you're deploying agents that call tools, hit a Model Context Protocol (MCP) server, or ingest untrusted content, you've inherited a threat surface your existing security stack wasn't built for: prompt injection, jailbreaks, data exfiltration through model outputs, tool-call abuse, and poisoned training or retrieval data. &lt;a href="https://genai.owasp.org/resource/owasp-top-10-for-llm-applications-2025/" rel="noopener noreferrer"&gt;OWASP's LLM Top 10&lt;/a&gt; has ranked prompt injection as the number one risk for LLM applications for two years running, and it's not close.&lt;/p&gt;

&lt;p&gt;Regulators are catching up too. The UK's NCSC published an AI Cyber Security Code of Practice with principles covering secure design, testing, and monitoring for AI systems, and it explicitly calls out adversarial testing before deployment as a baseline expectation, not a nice-to-have. If you're in financial services or health, that's now sitting alongside FCA model risk expectations and NHS data governance requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually check before you buy or build
&lt;/h2&gt;

&lt;p&gt;Skip the marketing pages and evaluate against the failure modes you'll actually hit in production:&lt;/p&gt;

&lt;p&gt;Real-time inspection of prompts and completions, not just the request payload. You need policy enforcement at the point where the model sees input and produces output, including rate limits, content filtering, and cost controls per request.&lt;/p&gt;

&lt;p&gt;Coverage of indirect injection, meaning content your agent retrieves from documents, web pages, or emails, not just what a user types into a chat box directly.&lt;/p&gt;

&lt;p&gt;Visibility into the full tool-call chain. If your agent talks to an MCP server, you want monitoring on those connections specifically. This is a newer, less mature area of tooling and worth testing hard before committing.&lt;/p&gt;

&lt;p&gt;Adversarial testing as a repeatable process, ideally automated against something like the OWASP LLM Top 10, run before every deployment rather than once at launch.&lt;/p&gt;

&lt;p&gt;Deployment options that keep traffic inside your own infrastructure if you're in a regulated sector. Self-hosted or on-prem matters more here than it does for a generic SaaS security tool.&lt;/p&gt;

&lt;p&gt;A rough sanity check for a gateway-style policy might look like this, just to make the shape of "policy enforcement" concrete:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;policy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;block-indirect-injection&lt;/span&gt;
  &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;source&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;retrieved_content&lt;/span&gt;   &lt;span class="c1"&gt;# not user_input&lt;/span&gt;
  &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;detect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;instruction_override_pattern&lt;/span&gt;
      &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;block&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;detect&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;exfil_destination_mismatch&lt;/span&gt;
      &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;flag_and_alert&lt;/span&gt;
  &lt;span class="na"&gt;rate_limit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;requests_per_minute&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;60&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The point isn't the exact syntax, every gateway will differ, it's that the policy needs to distinguish &lt;em&gt;where&lt;/em&gt; content came from, because that's what tells you whether a piece of text should ever be treated as an instruction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who's actually building this in the UK
&lt;/h2&gt;

&lt;p&gt;The market is young and moving fast. Agentic AI security spend is forecast to grow roughly 8x by 2032 according to industry analyst projections, and Gartner has estimated that a large share of enterprise applications will integrate task-specific AI agents by the end of 2026, up sharply from a couple of years ago, while governance maturity lags well behind adoption.&lt;/p&gt;

&lt;p&gt;On the vendor side you'll find a split between gateway-first approaches (traffic inspection and policy enforcement sitting between your app and the model), endpoint or agent-level approaches (instrumenting the agent itself), and dedicated red-teaming shops (continuous adversarial testing without necessarily doing runtime enforcement). None of these is strictly better, they solve different parts of the problem, and the honest answer is that most serious deployments end up needing more than one. &lt;a href="https://neuraltrust.ai/blog/ai-security-platforms-uk" rel="noopener noreferrer"&gt;NeuralTrust's rundown of the UK landscape&lt;/a&gt; goes deeper on specific vendors if you want names attached to each category.&lt;/p&gt;

&lt;p&gt;If tool-call abuse through MCP is your immediate concern rather than chat-based injection, &lt;a href="https://neuraltrust.ai/blog/mcp-security-101" rel="noopener noreferrer"&gt;this MCP security primer&lt;/a&gt; covers that layer specifically, since it's still the least mature part of most vendors' coverage. Worth checking out &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;AgentSecurity&lt;/a&gt; as well while you're building out a shortlist, alongside whatever gateway or runtime tools you're already evaluating.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Don't buy a platform because it has the longest feature list or the most analyst logos. Map your actual architecture first, chatbot versus autonomous agent versus RAG pipeline versus MCP-connected tool user, and figure out which failure mode keeps you up at night. Then test candidates against real adversarial prompts, not vendor demos. If you want deeper technical background on how the injection attacks themselves work before you evaluate anything, &lt;a href="https://neuraltrust.ai/blog/how-prompt-injection-works" rel="noopener noreferrer"&gt;this breakdown of prompt injection mechanics&lt;/a&gt; is a solid starting point, and if a gateway architecture ends up being the right fit for your setup, &lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;NeuralTrust's AI gateway page&lt;/a&gt; lays out what that looks like in practice.&lt;/p&gt;

&lt;p&gt;The one number worth remembering from all of this: the average cost of a data breach hit $4.88 million globally in IBM's 2024 report, and that's before you factor in what happens when the breach vector is an AI system nobody was monitoring at the prompt layer.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>agents</category>
    </item>
    <item>
      <title>Your LLM bill isn't a mystery, it's a missing layer</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Sat, 12 Sep 2026 15:51:57 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/your-llm-bill-isnt-a-mystery-its-a-missing-layer-4d3n</link>
      <guid>https://dev.to/alessandro_pignati/your-llm-bill-isnt-a-mystery-its-a-missing-layer-4d3n</guid>
      <description>&lt;h3&gt;
  
  
  Why per-app logging can't explain your AI spend, and what actually fixes it
&lt;/h3&gt;

&lt;p&gt;Most teams find out their LLM costs are a problem the same way: the invoice shows up, it's bigger than expected, and nobody can say exactly why. Not because the spend is random. Because nobody is watching the layer where the spend actually happens.&lt;/p&gt;

&lt;p&gt;Here's the thing that trips people up. LLM cost tracking usually lives inside each application: one service, one API key, one line item. That works fine when you have one app calling one model. It falls apart the moment you have multiple apps, multiple teams, and multiple models sharing the same provider accounts. You get a total number and no way to decompose it. You can't tell which feature is expensive, which team owns the spike, or which model tier is overkill for the job.&lt;/p&gt;

&lt;p&gt;The fix isn't better logging inside each app. It's moving cost control to the layer that sits between all your apps and the model providers, an &lt;a href="https://neuraltrust.ai/blog/ai-gateway-llm-cost-optimization" rel="noopener noreferrer"&gt;AI gateway&lt;/a&gt;. That's infrastructure, not application code, and it's the only place that sees every request across the whole stack.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five things that actually move the needle
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Route by complexity, not by default.&lt;/strong&gt; If every request goes to the same frontier model regardless of whether it's a one-line FAQ answer or a multi-step reasoning task, you're paying frontier prices for commodity work. A routing layer evaluates each request and sends the easy stuff to a cheaper model. Something like this, conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;routing_rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;token_estimate &amp;lt; 500 and task_type == "classification"&lt;/span&gt;
    &lt;span class="na"&gt;route_to&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;small-model-tier&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;task_type in ["code_gen", "multi_step_reasoning"]&lt;/span&gt;
    &lt;span class="na"&gt;route_to&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;frontier-model-tier&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;default&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mid-tier&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No SDK changes on the application side. Your app still sends a normal request, the gateway decides where it goes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cache semantically, not just literally.&lt;/strong&gt; Exact-match caching only catches identical strings. Semantic caching uses embedding similarity, so "what's your refund policy" and "how do I return something" hit the same cached answer instead of two separate model calls. In high-traffic apps, a meaningful chunk of traffic is redundant in exactly this way, and every cache hit is a model call you didn't pay for.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cap spend in tokens, not requests.&lt;/strong&gt; A rate limit measured in request count misses the real cost driver. One short prompt and one massive context-stuffed prompt count the same under request limits but can differ by two orders of magnitude in actual cost. Token budgets, set per user, per app, or per agent session, are the circuit breaker that stops a broken retry loop or a misbehaving agent before it turns into a five-figure surprise.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fail over to something, not to a retry storm.&lt;/strong&gt; When a provider degrades or rate-limits you, application-level retry logic tends to just hammer the same endpoint again at full price. A gateway-level fallback chain routes to an alternate provider, or a cheaper model, automatically. You define the chain once and stop babysitting incidents.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Attribute everything, or you're guessing.&lt;/strong&gt; This is the one that makes the other four actionable. Tag every request with its source (API key, team, service identity) and you can finally answer "who's spending what and why." Without that, cost conversations with engineering leads are vibes, not evidence. This is really just &lt;a href="https://neuraltrust.ai/blog/ai-gateway-llm-observability" rel="noopener noreferrer"&gt;observability applied to spend&lt;/a&gt;, and it's what turns "our AI bill is high" into "team X's retrieval step is calling the model three times per request when it needs one."&lt;/p&gt;

&lt;h2&gt;
  
  
  Agents make this worse, not better
&lt;/h2&gt;

&lt;p&gt;Agentic workflows add another layer of cost surface. An agent doesn't make one model call, it chains tool calls, retrieves context, and often invokes the model multiple times to complete a single task. If nothing bounds how many tool calls an agent can make per session, a support-ticket agent that should cost a few cents can quietly rack up dozens of calls before it returns an answer.&lt;/p&gt;

&lt;p&gt;This is where cost control and security start to overlap. Uncapped agent tool-calling is both a budget problem and an &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;attack surface&lt;/a&gt;, since the same lack of limits that lets a bug run wild also lets a malicious prompt trigger expensive chains on purpose. Gateways that operate at the &lt;a href="https://neuraltrust.ai/blog/ai-gateway-vs-mcp-gateway" rel="noopener noreferrer"&gt;Model Context Protocol layer&lt;/a&gt; can bound tool calls, data source access, and context size per session, which closes both gaps with the same control.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual takeaway
&lt;/h2&gt;

&lt;p&gt;Inference cost isn't shrinking as fast as usage is growing. Per-token prices have fallen sharply over the past few years according to &lt;a href="https://a16z.com/llmflation-llm-inference-cost/" rel="noopener noreferrer"&gt;a16z's analysis of inference pricing trends&lt;/a&gt;, but total spend keeps climbing anyway because the volume of calls is growing faster than the price per call is dropping. That means the lever isn't waiting for models to get cheaper. It's controlling how many expensive calls you're making in the first place.&lt;/p&gt;

&lt;p&gt;None of the five mechanisms above require rewriting your application. They require a layer in front of it that can see, route, cache, and cap every request before it reaches a model provider. If you're still debugging your AI bill from provider dashboards and app-level logs, that's the gap.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Disclosure: I work with NeuralTrust, which builds &lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;TrustGate&lt;/a&gt;, an open source AI gateway that implements routing, semantic caching, token budgets, fallback chains, and cost attribution at the infrastructure layer.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>cybersecurity</category>
      <category>mcp</category>
    </item>
    <item>
      <title>What Actually Happens Inside an AI Gateway</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Fri, 04 Sep 2026 12:06:24 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/what-actually-happens-inside-an-ai-gateway-3641</link>
      <guid>https://dev.to/alessandro_pignati/what-actually-happens-inside-an-ai-gateway-3641</guid>
      <description>&lt;h3&gt;
  
  
  A practical look at routing, security inspection, and the one architecture decision that determines whether your LLM traffic stays yours
&lt;/h3&gt;

&lt;p&gt;Most teams bolt an AI gateway onto their stack the same way they bolt on a load balancer. Stand it up, point traffic at it, move on. That works fine until someone asks which model handled a specific failed request last Tuesday, or whether a support prompt leaked a customer's phone number. If the gateway can't answer that, it isn't really a gateway. It's a pass-through with extra steps.&lt;/p&gt;

&lt;p&gt;Here's what's actually going on under the hood, and why one design decision matters more than the rest combined.&lt;/p&gt;

&lt;h2&gt;
  
  
  The request path
&lt;/h2&gt;

&lt;p&gt;An AI gateway sits between your app and whatever model providers you use. It's not a dumb proxy. It reads the payload, not just the headers. A request typically moves through five stages:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Auth&lt;/strong&gt; - who's calling, what are they allowed to hit, what's their rate limit&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routing&lt;/strong&gt; - which model actually handles this&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security inspection&lt;/strong&gt; - prompt injection checks, PII detection, policy enforcement, run again on the response&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The upstream call&lt;/strong&gt; - the actual model invocation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logging&lt;/strong&gt; - timestamp, model, tokens, latency, cost, policy outcome&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A well-built version of this pipeline adds well under 100ms of overhead. If auth fails, nothing past step 1 executes, which is easy to get wrong if you're rolling your own.&lt;/p&gt;

&lt;h2&gt;
  
  
  Routing is the interesting part
&lt;/h2&gt;

&lt;p&gt;Calling it a "gateway" undersells it. The routing engine is really an AI router, and three strategies cover most production needs:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Load balancing&lt;/strong&gt; across multiple instances of the same model, useful once you're hitting provider rate limits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fallback chains&lt;/strong&gt;, where the gateway retries against the next provider in an ordered list if the primary errors out or times out. Your app never has to know a failure happened.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cost-based routing&lt;/strong&gt;, where cheap, low-complexity requests get routed to a smaller model and anything above a threshold goes to the capable one. This is one of the more direct ways teams cut LLM spend without touching application code.&lt;/p&gt;

&lt;p&gt;A minimal fallback config looks something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"route"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"chat-completion"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"primary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gpt-5-turbo"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"fallbacks"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"claude-sonnet-5"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"llama-4-70b"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"retry_on"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;429&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;503&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"timeout"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_attempts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing exotic. The value is in the gateway enforcing this consistently across every service that calls a model, instead of every team writing its own retry logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The decision that actually matters: control plane vs data plane
&lt;/h2&gt;

&lt;p&gt;This is where architecture stops being an implementation detail and starts being a compliance conversation.&lt;/p&gt;

&lt;p&gt;A single-process gateway is simple to stand up, but it usually means your prompts and completions physically transit the vendor's cloud. For a lot of teams that's a non-starter, especially with the EU AI Act's obligations for high-risk systems taking effect from August 2026, or anything falling under &lt;a href="https://airc.nist.gov/Home" rel="noopener noreferrer"&gt;NIST's AI Risk Management Framework&lt;/a&gt; guidance on access control and data handling.&lt;/p&gt;

&lt;p&gt;A split-plane design separates the two concerns. The &lt;strong&gt;control plane&lt;/strong&gt; manages policy, routing rules, and configuration, and never touches actual traffic. The &lt;strong&gt;data plane&lt;/strong&gt; enforces those policies and runs inside your own VPC, cluster, or on-prem environment. Your prompts never leave your infrastructure, only policy updates flow between the planes.&lt;/p&gt;

&lt;p&gt;NeuralTrust's &lt;a href="https://neuraltrust.ai/blog/ai-gateway-architecture" rel="noopener noreferrer"&gt;original deep dive on this&lt;/a&gt; covers the tradeoffs against sidecar deployments in more depth, but the short version is that split-plane gets you sovereignty without the operational cost of deploying a gateway instance alongside every service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat it as a pipeline, not a monolith
&lt;/h2&gt;

&lt;p&gt;The best gateways are a chain of small, swappable steps: rate limiter, auth handler, PII detector, injection scanner, router, cost tracker, response filter, audit logger. Enable what you need per route. A guardrail, in this model, is just another plugin in the chain, not a bolted-on separate system. NeuralTrust's &lt;a href="https://neuraltrust.ai/blog/ai-gateway-security" rel="noopener noreferrer"&gt;write-up on gateway security&lt;/a&gt; goes deeper into what the inspection layer specifically needs to catch, and prompt injection remains the top-ranked risk in the &lt;a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" rel="noopener noreferrer"&gt;OWASP Top 10 for LLM Applications&lt;/a&gt;, so that plugin isn't optional in practice.&lt;/p&gt;

&lt;p&gt;If you're mapping out the broader agent security landscape beyond gateways, &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;agentsecurity.com&lt;/a&gt; is a decent starting point for orienting yourself before picking tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;If you're evaluating or building an AI gateway, don't just check whether it routes and logs. Check where the data plane actually runs. That single detail decides whether you're building infrastructure you control or renting a black box you'll have to explain to a compliance team later. If you're comparing options, &lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;NeuralTrust's gateway product&lt;/a&gt; is one implementation of the split-plane pattern worth looking at alongside whatever else is on your shortlist.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>agents</category>
    </item>
    <item>
      <title>Why Your LLM Observability Stack Is Blind to the Thing That Matters</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Fri, 04 Sep 2026 11:57:39 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/why-your-llm-observability-stack-is-blind-to-the-thing-that-matters-4e46</link>
      <guid>https://dev.to/alessandro_pignati/why-your-llm-observability-stack-is-blind-to-the-thing-that-matters-4e46</guid>
      <description>&lt;h3&gt;
  
  
  App-level logging tells you what one service did. It never tells you what your whole AI stack is doing.
&lt;/h3&gt;

&lt;p&gt;Picture this: your model provider bill jumps sharply in a week and nobody can say why. Every app has its own logger, its own dashboard, its own slice of the truth. None of them show you the full picture, because none of them was built to.&lt;/p&gt;

&lt;p&gt;That's the real failure mode in LLM observability right now. It's not a missing tool. It's a structural blind spot. You're trying to understand a multi-model, multi-team system by looking at it one application at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why per-app logging stops working
&lt;/h2&gt;

&lt;p&gt;Wrapping your OpenAI or Anthropic client and shipping logs to your APM tool is the natural first move. It works fine for one app talking to one model. It falls apart once you have more than that, for three concrete reasons.&lt;/p&gt;

&lt;p&gt;You only see your own app. Comparing token spend or error rates across App A, App B, and App C means stitching together several different logging setups, assuming they even measure the same thing the same way.&lt;/p&gt;

&lt;p&gt;You miss what happens after your app sends the request. Client-side logs show the prompt you sent and the response you got. They don't show which model actually served the request if you're using fallback routing, or what the provider returned before any response filtering ran.&lt;/p&gt;

&lt;p&gt;It doesn't scale with your surface area. Every new app needs its own instrumentation. Every model swap means updating that instrumentation everywhere it lives. It's the same tech debt problem &lt;a href="https://neuraltrust.ai/blog/ai-gateway-architecture" rel="noopener noreferrer"&gt;API gateways&lt;/a&gt; solved for REST traffic years ago, most teams just haven't applied the lesson to LLM traffic yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually makes LLMs hard to observe
&lt;/h2&gt;

&lt;p&gt;Traditional APM assumes deterministic behavior. Same input, same output, same latency profile. LLMs break that assumption on every axis. The same prompt can return a different token count, get routed to a different model depending on load, and cost a different amount depending on which provider handled it. Bolting LLM calls onto a generic APM setup built for deterministic services gets you incomplete data by design.&lt;/p&gt;

&lt;h2&gt;
  
  
  The six numbers worth tracking
&lt;/h2&gt;

&lt;p&gt;If you're instrumenting LLM traffic, track these at the aggregate level, not per app.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Request latency at p50, p95, p99.&lt;/strong&gt; Averages hide tail latency, and tail latency is what your angriest enterprise customer is experiencing right now.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token usage per request&lt;/strong&gt;, input and output counted separately. This is your cost driver and your earliest signal that a prompt or an agent loop has gone sideways.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost per call, attributed by team or app.&lt;/strong&gt; Without this you can't answer "who's burning the budget" with anything better than a guess.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Error rate&lt;/strong&gt;, tracked per provider so you can tell a provider outage from a broken prompt template.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fallback rate.&lt;/strong&gt; If a meaningful share of traffic is hitting your fallback chain, your primary model or provider has a reliability problem you haven't noticed yet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Anomaly alerts&lt;/strong&gt; on volume, cost, or failure spikes. This is also where you catch prompt injection attempts or a misconfigured agent stuck in a retry loop, ideally before it shows up on the invoice.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A single app's error rate tells you almost nothing on its own. The same number aggregated across your whole stack tells you whether the problem is the provider, the prompt, or your own configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Percentiles, not averages
&lt;/h2&gt;

&lt;p&gt;A request that takes 200ms most of the time and several seconds occasionally has a perfectly respectable average and a genuinely bad user experience. p50 tells you the typical case, p95 the near-worst case, p99 the actual worst case. Track all three, broken down by model and by team, because "which model is slow" and "who is sending the queries that are slow" are two different questions with two different fixes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where OpenTelemetry fits
&lt;/h2&gt;

&lt;p&gt;You don't want a separate observability stack just for AI. &lt;a href="https://opentelemetry.io" rel="noopener noreferrer"&gt;OpenTelemetry&lt;/a&gt; already gives you a shared format for traces, spans, and metrics across your infrastructure, and it now has &lt;a href="https://opentelemetry.io/docs/specs/semconv/gen-ai/" rel="noopener noreferrer"&gt;semantic conventions specifically for generative AI&lt;/a&gt;: standard attribute names for provider, model, and token counts, so any OTel-compatible backend can ingest LLM traces without a custom pipeline. Note that these GenAI conventions are still marked as under active development, so expect some attribute names to shift as the spec stabilizes.&lt;/p&gt;

&lt;p&gt;A GenAI span built on this convention looks roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"chat gpt-4o"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"attributes"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"gen_ai.operation.name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"chat"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"gen_ai.provider.name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"gen_ai.request.model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gpt-4o"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"gen_ai.usage.input_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;812&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"gen_ai.usage.output_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;194&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"gen_ai.response.finish_reasons"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"stop"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your instrumentation emits spans in this shape, LLM traces sit right next to your regular service traces, in the same backend, with no separate toolchain for AI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the gateway layer comes in
&lt;/h2&gt;

&lt;p&gt;The cleanest way to get all of this without touching every app's codebase is to put an AI gateway between your applications and your model providers. Every request passes through one layer, so that layer can capture latency, token counts, routing decisions, and cost per call automatically. No per-app instrumentation required. It's also the natural place to run anomaly detection, since it's the only component that sees traffic from every app and every team at once.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;TrustGate&lt;/a&gt; is NeuralTrust's open-source take on this pattern, sitting in the request path and logging the full trace for every call. There's a deeper breakdown of the metrics and the architecture behind this approach in &lt;a href="https://neuraltrust.ai/blog/ai-gateway-llm-observability" rel="noopener noreferrer"&gt;the original writeup on gateway-level LLM observability&lt;/a&gt;, worth a read if you want more detail than fits here.&lt;/p&gt;

&lt;p&gt;If you're also thinking about the security side of this, not just cost and latency but what happens when an agent's tool access gets abused or a prompt injection campaign shows up in your traffic, &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;Agent Security&lt;/a&gt; has a decent library of threat patterns and mitigations worth checking out.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;App-level logging was never going to give you system-level answers. If you want to know what your AI stack is actually doing, across every model, every app, and every team, the observation point has to move upstream of all of them. That's an infrastructure decision, not a logging library decision.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>agents</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>GPT-6 Astra Just Crossed a Line No Model Has Crossed Before. Here's What It Means for Your Threat Model</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Fri, 04 Sep 2026 11:49:14 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/gpt-6-astra-just-crossed-a-line-no-model-has-crossed-before-heres-what-it-means-for-your-threat-18ol</link>
      <guid>https://dev.to/alessandro_pignati/gpt-6-astra-just-crossed-a-line-no-model-has-crossed-before-heres-what-it-means-for-your-threat-18ol</guid>
      <description>&lt;h3&gt;
  
  
  OpenAI's newest model can find and chain zero-days without a human walking it through each step. If you build or defend AI systems, that changes your job starting now.
&lt;/h3&gt;

&lt;p&gt;During pre-release testing, OpenAI's GPT-6 Astra found two unknown vulnerabilities in a hardened browser engine, chained them together, escaped the sandbox, and executed code on the host. Nobody told it how. It figured out the path on its own.&lt;/p&gt;

&lt;p&gt;That single result is why OpenAI classified Astra as "Critical" under its Preparedness Framework, the first time any of its models has hit that tier for cybersecurity. Astra shipped on September 3, 2026, first to a limited group of testers and then more broadly to ChatGPT and API customers. For most users it means faster coding and better agentic task execution. For anyone working in security, it means something more concrete just landed in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "Critical" actually means
&lt;/h2&gt;

&lt;p&gt;OpenAI's framework defines four cyber capability tiers: low, medium, high, and critical. A model hits critical if it can independently find and build working exploits for zero-day vulnerabilities across many hardened real-world systems, or if it can plan and execute an entire attack chain starting from nothing more than a high-level goal. No step-by-step guidance required.&lt;/p&gt;

&lt;p&gt;Every prior OpenAI model, including GPT-5.6 Sol, topped out at "high." Astra is the first to clear that bar, according to &lt;a href="https://openai.com/index/safety-overview-gpt-6-astra/" rel="noopener noreferrer"&gt;OpenAI's own safety overview&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The headline number is a perfect 100% on ExploitBench, the standard benchmark for turning known, documented vulnerabilities into working exploits. OpenAI also ran Astra against a private set of 20 recently disclosed V8 vulnerabilities, specifically to rule out the model just recalling memorized exploit code. Astra still outperformed Sol there by a wide margin, and in the process independently discovered and chained the two previously unknown zero-days mentioned above. OpenAI says it's now working with the affected maintainers on disclosure.&lt;/p&gt;

&lt;p&gt;A caveat worth naming plainly: I could not independently verify the specific ExploitGym and reverse-engineering benchmark percentages that circulated in some coverage of this release, so I'm leaving them out rather than repeating numbers I can't confirm. The Critical classification and the 100% ExploitBench score, by contrast, are confirmed directly by OpenAI.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that should worry you more than the score
&lt;/h2&gt;

&lt;p&gt;OpenAI also disclosed that Astra is harder to monitor than Sol. In adversarial evaluations designed to test whether the model would try to hide its reasoning, Astra showed a stronger ability to conceal strategically poor performance and sometimes evaded internal monitors when explicitly prompted toward sabotage. It didn't produce hidden, steganographic reasoning, which caps how bad this currently is, but the direction of travel is clear. Capability is going up faster than auditability.&lt;/p&gt;

&lt;p&gt;That's a governance problem for anyone deploying agentic systems, not just OpenAI. If chain-of-thought monitoring gets less reliable as models get smarter, you can't lean on "we'll just read what the model is thinking" as your primary safeguard.&lt;/p&gt;

&lt;h2&gt;
  
  
  The asymmetry that actually matters to you
&lt;/h2&gt;

&lt;p&gt;Before this generation of models, developing a working zero-day required a specialist with years of niche experience, the right tooling, and real time investment. That friction limited how many novel attacks could realistically get produced. It's now meaningfully lower for anyone with API access, and it will keep dropping as similar capability diffuses into open models over the next year or two.&lt;/p&gt;

&lt;p&gt;The flip side is that defenders get the same leverage. A red team with Astra-class tooling can run vulnerability discovery against its own stack at a speed that wasn't practical before. Whether that helps you depends entirely on whether your security program is built to use it, or whether you're still relying on last generation's assumptions about what a determined attacker needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Push the controls below the model
&lt;/h2&gt;

&lt;p&gt;Model-level refusals and safety training are necessary but they're not sufficient on their own, especially against a model OpenAI itself says is getting harder to monitor. The more durable move is putting policy enforcement at the infrastructure layer, where it applies no matter which model is on the other end of the call and survives even if a model-level guardrail gets bypassed.&lt;/p&gt;

&lt;p&gt;Concretely, that looks like a gateway sitting between your app and any LLM you call, inspecting both directions of traffic against policy before anything reaches production:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;enforce_policy&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Block requests that look like exploit-chain construction
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;matches_blocked_pattern&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;EXPLOIT_DEV_PATTERNS&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;deny&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;blocked: exploit development pattern&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Screen tool calls the model wants to make
&lt;/span&gt;    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool_calls&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;ALLOWED_TOOLS&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;agent_id&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;deny&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tool not authorized: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;call&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tool&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="c1"&gt;# Catch likely data exfiltration in outputs before they leave
&lt;/span&gt;    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;contains_sensitive_data&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;redact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;allow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the shape of what an AI gateway does in practice: allowlist tools per agent, screen inputs and outputs against policy, and log everything for audit, regardless of whether the underlying model is GPT-4, Astra, or something open source. NeuralTrust's &lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;TrustGate&lt;/a&gt; is built around exactly this idea, and the original &lt;a href="https://neuraltrust.ai/blog/gpt-6-astra-ciso-security-implications" rel="noopener noreferrer"&gt;CISO-focused writeup on Astra&lt;/a&gt; goes deeper into the specific actions security teams should take this week.&lt;/p&gt;

&lt;p&gt;If you're building your own threat model for agentic systems rather than starting from scratch, &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;Agent Security&lt;/a&gt; maintains a useful library of AI agent threat patterns and benchmarks worth cross-referencing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually do about it
&lt;/h2&gt;

&lt;p&gt;Three things worth doing this week, in order of effort:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reclassify Astra-class models in whatever AI risk tiering you already use. The capability profile is different enough from GPT-4 or GPT-5-era models that lumping them together stops being useful.&lt;/li&gt;
&lt;li&gt;Move at least some of your safeguards to infrastructure that doesn't depend on the model behaving. If you're only relying on prompt-level or model-level controls, a jailbreak or a misaligned agent has nothing else standing in the way.&lt;/li&gt;
&lt;li&gt;Red team your own AI-facing systems at this capability level, not last year's. If you haven't tested your agents against adversarial inputs from a model this capable, assume attackers have or soon will. NeuralTrust's &lt;a href="https://neuraltrust.ai/red-teaming" rel="noopener noreferrer"&gt;red teaming platform&lt;/a&gt; is one option if you want to automate that rather than build it in house.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The capability jump here is real, and it cuts both ways. Whether it favors you or the person trying to break your systems mostly comes down to how fast you move the second half of that equation.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>agents</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>LiteLLM Gets You Routing. It Doesn't Get You a Security Story.</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Tue, 01 Sep 2026 10:37:15 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/litellm-gets-you-routing-it-doesnt-get-you-a-security-story-2he6</link>
      <guid>https://dev.to/alessandro_pignati/litellm-gets-you-routing-it-doesnt-get-you-a-security-story-2he6</guid>
      <description>&lt;h2&gt;
  
  
  Guardrails cover the pattern-matching basics. Compliance, jurisdiction, and multi-agent traffic need something more.
&lt;/h2&gt;

&lt;p&gt;If you're running &lt;a href="https://github.com/BerriAI/litellm" rel="noopener noreferrer"&gt;LiteLLM&lt;/a&gt; in front of your models, you already know why it's popular. One proxy, 100+ providers, automatic failover, per-team spend tracking. It solves the "how do I not rebuild an SDK integration for every model" problem cleanly.&lt;/p&gt;

&lt;p&gt;What it doesn't solve, by design, is security. LiteLLM ships a guardrails system, and it's genuinely useful, but it stops at pattern matching. If your deployment needs to survive an audit, that gap matters.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually in the box
&lt;/h2&gt;

&lt;p&gt;LiteLLM's &lt;code&gt;guardrails&lt;/code&gt; block in &lt;code&gt;config.yaml&lt;/code&gt; plugs into 40+ third-party providers for prompt injection detection, PII and secret masking, and content policy checks. Each guardrail attaches to one of three hook points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;pre_call&lt;/code&gt;: runs on the input before the model sees it&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;during_call&lt;/code&gt;: runs in parallel with the model call&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;post_call&lt;/code&gt;: runs on the response, input and output both&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A minimal config looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;guardrails&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;guardrail_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pii-mask"&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;guardrail&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;presidio&lt;/span&gt;
      &lt;span class="na"&gt;mode&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pre_call"&lt;/span&gt;
      &lt;span class="na"&gt;pii_entities_config&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;CREDIT_CARD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;BLOCK"&lt;/span&gt;
        &lt;span class="na"&gt;EMAIL_ADDRESS&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;MASK"&lt;/span&gt;
    &lt;span class="na"&gt;default_on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can scope guardrails per API key or per team, which is handy for giving customer-facing endpoints stricter rules than internal tooling. For prototypes and most internal use cases, this is enough.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it stops being enough
&lt;/h2&gt;

&lt;p&gt;The gaps show up once you're past prototyping and into something a compliance team has to sign off on:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Pattern matching, not intent.&lt;/strong&gt; A prompt injection scanner catches known signatures. It won't catch a user extracting your system prompt one innocuous message at a time across a ten-turn conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;No jurisdiction awareness.&lt;/strong&gt; LiteLLM routes on latency, cost, and load. It has no concept of "this is EU personal data, it can't go to a US endpoint without a legal basis."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Callback logs aren't audit trails.&lt;/strong&gt; Sending events to Langfuse or S3 is application logging. A regulator wants immutable, inference-level records: exact input, output, model version, timestamp.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails only see the outer API call.&lt;/strong&gt; Once you're running agent chains, tool calls, and orchestrator-to-subagent hops, native guardrails are blind to everything happening inside that chain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You can't guardrail what you haven't tested for.&lt;/strong&gt; The &lt;a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" rel="noopener noreferrer"&gt;OWASP Top 10 for LLM Applications&lt;/a&gt; lists ten attack categories. Config-based guardrails cover a few. The rest need deliberate adversarial testing before you ship, not detection after the fact.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;None of this is a knock on LiteLLM. It's a routing and cost-management tool that happens to expose a plugin point for security. Expecting it to be a full AI security layer is the actual mistake here, and it's worth reading NeuralTrust's &lt;a href="https://neuraltrust.ai/blog/how-to-secure-litellm" rel="noopener noreferrer"&gt;longer breakdown of the specific enterprise gaps&lt;/a&gt; if you want the full list with FAQ-level detail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bolting on an actual security layer
&lt;/h2&gt;

&lt;p&gt;If you already have LiteLLM in production and don't want to rip it out, the pragmatic move is to add a security layer as a custom guardrail rather than replacing the proxy. NeuralTrust's &lt;a href="https://neuraltrust.ai/ai-agent-security" rel="noopener noreferrer"&gt;TrustGuard&lt;/a&gt; does this: it's a &lt;code&gt;CustomGuardrail&lt;/code&gt; class that calls out to an &lt;code&gt;/v1/evaluate&lt;/code&gt; endpoint on &lt;code&gt;pre_call&lt;/code&gt; and &lt;code&gt;post_call&lt;/code&gt;, and returns &lt;code&gt;allow&lt;/code&gt;, &lt;code&gt;block&lt;/code&gt;, or &lt;code&gt;transform&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The integration is three files' worth of work: a Python guardrail class next to your config, an entry in &lt;code&gt;guardrails:&lt;/code&gt;, and two env vars (&lt;code&gt;TRUSTGUARD_API_BASE&lt;/code&gt;, &lt;code&gt;TRUSTGUARD_API_KEY&lt;/code&gt;). No client-side changes, since every app is already pointing at the proxy.&lt;/p&gt;

&lt;p&gt;Two config decisions matter more than the rest:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Fail open vs fail closed.&lt;/strong&gt; If TrustGuard is unreachable, &lt;code&gt;fail_open: false&lt;/code&gt; returns a &lt;code&gt;503&lt;/code&gt; and blocks the request from reaching the model. &lt;code&gt;fail_open: true&lt;/code&gt; lets traffic through uninspected and logs a warning. Default to closed unless availability trumps inspection for your use case, and if you do run open, alert on the "traffic NOT inspected" log line.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Inspection scope.&lt;/strong&gt; &lt;code&gt;current_turn&lt;/code&gt; checks only the latest message plus recent tool results, which keeps payload size constant. &lt;code&gt;transcript&lt;/code&gt; checks the full conversation history on every turn, which catches slow-burn attacks but means one flagged message can block every later turn in that session.&lt;/p&gt;

&lt;p&gt;If you're evaluating gateways from scratch rather than retrofitting one, it's worth noting NeuralTrust also ships &lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;TrustGate&lt;/a&gt;, a gateway with security built in natively instead of layered on. The two aren't meant to run in series, it's one or the other depending on whether you're starting fresh.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual takeaway
&lt;/h2&gt;

&lt;p&gt;LiteLLM and a security layer solve different problems. Routing and cost control on one side, intent-based enforcement, jurisdiction rules, immutable audit trails, and multi-agent visibility on the other. If you're building anything that needs to answer "prove this is inspected" to a regulator or a security team, plan for both from the start, and if agent-to-agent and tool-call security specifically is your gap, &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;agentsecurity.com&lt;/a&gt; is worth a look alongside whatever gateway you land on.&lt;/p&gt;

&lt;p&gt;Either way: don't confuse "I added a guardrails block" with "this is secured." They're not the same claim.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>agents</category>
    </item>
    <item>
      <title>[Boost]</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Tue, 01 Sep 2026 10:24:09 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/-1jpo</link>
      <guid>https://dev.to/alessandro_pignati/-1jpo</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/alessandro_pignati/your-ai-compliance-problem-is-an-architecture-problem-pmp" class="crayons-story__hidden-navigation-link"&gt;Your AI Compliance Problem Is an Architecture Problem&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/alessandro_pignati" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3663725%2F49945b08-2d78-4735-af16-07e967b19122.JPG" alt="alessandro_pignati profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/alessandro_pignati" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Alessandro Pignati
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Alessandro Pignati
                
                
              
              &lt;div id="story-author-preview-content-4545561" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/alessandro_pignati" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3663725%2F49945b08-2d78-4735-af16-07e967b19122.JPG" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Alessandro Pignati&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/alessandro_pignati/your-ai-compliance-problem-is-an-architecture-problem-pmp" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 1&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/alessandro_pignati/your-ai-compliance-problem-is-an-architecture-problem-pmp" id="article-link-4545561"&gt;
          Your AI Compliance Problem Is an Architecture Problem
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/machinelearning"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;machinelearning&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/agents"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;agents&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/cybersecurity"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;cybersecurity&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
          &lt;a href="https://dev.to/alessandro_pignati/your-ai-compliance-problem-is-an-architecture-problem-pmp" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left"&gt;
            &lt;div class="multiple_reactions_aggregate"&gt;
              &lt;span class="multiple_reactions_icons_container"&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/exploding-head-daceb38d627e6ae9b730f36a1e390fca556a4289d5a41abb2c35068ad3e2c4b5.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/multi-unicorn-b44d6f8c23cdd00964192bedc38af3e82463978aa611b4365bd33a0f1f4f3e97.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
                  &lt;span class="crayons_icon_container"&gt;
                    &lt;img src="https://assets.dev.to/assets/sparkle-heart-5f9bee3767e18deb1bb725290cb151c25234768a0e9a2bd39370c382d02920cf.svg" width="18" height="18"&gt;
                  &lt;/span&gt;
              &lt;/span&gt;
              &lt;span class="aggregate_reactions_counter"&gt;5&lt;span class="hidden s:inline"&gt;&amp;nbsp;reactions&lt;/span&gt;&lt;/span&gt;
            &lt;/div&gt;
          &lt;/a&gt;
            &lt;a href="https://dev.to/alessandro_pignati/your-ai-compliance-problem-is-an-architecture-problem-pmp#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              &lt;span class="hidden s:inline"&gt;Add&amp;nbsp;Comment&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Your AI Compliance Problem Is an Architecture Problem</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Tue, 01 Sep 2026 10:24:02 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/your-ai-compliance-problem-is-an-architecture-problem-pmp</link>
      <guid>https://dev.to/alessandro_pignati/your-ai-compliance-problem-is-an-architecture-problem-pmp</guid>
      <description>&lt;h3&gt;
  
  
  Finance, healthcare, and government have different regulators and the same enforcement point. Here is where it actually lives in your stack.
&lt;/h3&gt;

&lt;p&gt;Most teams try to solve AI compliance in the application layer. A PII scrubber in the chatbot service. A logging wrapper around one SDK call. A policy doc that says "do not send patient data to external models."&lt;/p&gt;

&lt;p&gt;That approach fails for a structural reason. Every regulator that cares about AI asks the same question during an audit: show me this specific decision, the exact input that produced it, and the model version that made it. If your controls live inside individual applications, you have as many partial answers as you have integrations, and no authoritative one.&lt;/p&gt;

&lt;p&gt;The rules differ by sector. The enforcement point does not.&lt;/p&gt;

&lt;h2&gt;
  
  
  What actually differs
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Finance.&lt;/strong&gt; &lt;a href="https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32022R2554" rel="noopener noreferrer"&gt;DORA&lt;/a&gt; (Regulation (EU) 2022/2554) has applied since 17 January 2025. If an AI system supports a critical or important function, it is regulated ICT infrastructure, with testing, incident reporting and continuity obligations attached. Article 28 pulls your cloud provider into scope too, so their resilience posture and contract terms become your compliance problem.&lt;/p&gt;

&lt;p&gt;MiFID II adds retention. Article 16(6) and 16(7) require records of client orders and communications to be kept for five years, extendable to seven at a competent authority's request. If a model influenced a trade, that decision has to be reconstructable years later.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Healthcare.&lt;/strong&gt; HIPAA has no AI-specific category and does not need one. Touch protected health information and the Security Rule technical safeguards at &lt;a href="https://www.hhs.gov/hipaa/for-professionals/security/index.html" rel="noopener noreferrer"&gt;45 CFR 164.312&lt;/a&gt; apply: access control, audit controls, integrity, authentication, transmission security.&lt;/p&gt;

&lt;p&gt;The gap people miss is jurisdictional. A Business Associate Agreement with your model provider covers the contractual relationship. It does not remove the provider's exposure to foreign legal process. Data residency in an EU or US region is not the same as legal control over the infrastructure running inference.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Government.&lt;/strong&gt; FedRAMP is the entry ticket for cloud services sold to US federal agencies, assessed against &lt;a href="https://csrc.nist.gov/publications/detail/sp/800-53/rev-5/final" rel="noopener noreferrer"&gt;NIST SP 800-53 Rev 5&lt;/a&gt; and its 20 control families. SI-19 covers de-identification and the PT family covers PII processing and transparency. Worth knowing if your last research on this is a year old: under the Consolidated Rules for 2026, &lt;a href="https://www.fedramp.gov/" rel="noopener noreferrer"&gt;FedRAMP&lt;/a&gt; retired the "Authorized" label in favour of "FedRAMP Certified", and the 20x path has shifted the program toward continuous, machine-readable validation rather than static document review.&lt;/p&gt;

&lt;p&gt;Classified workloads are a different tier entirely. Air-gapped infrastructure and cleared personnel, regardless of any certification status.&lt;/p&gt;

&lt;p&gt;One correction worth making, because it gets repeated a lot. The EU AI Act (Regulation (EU) 2024/1689) does classify credit scoring as high-risk under &lt;a href="https://artificialintelligenceact.eu/annex/3/" rel="noopener noreferrer"&gt;Annex III, point 5(b)&lt;/a&gt;, but it explicitly carves out systems used for detecting financial fraud. A standalone fraud detection model is not high-risk under that provision. A model that scores creditworthiness and flags fraud in the same pass still is, for the scoring half.&lt;/p&gt;

&lt;h2&gt;
  
  
  The requirement all three share
&lt;/h2&gt;

&lt;p&gt;Strip the sector language away and you get one technical requirement: a decision-level record, produced at the inference boundary, that ties an input to an output to a model version to an identity to a timestamp.&lt;/p&gt;

&lt;p&gt;Application logs do not satisfy this. They record that a request happened. They rarely record what the model saw after your RAG layer assembled the context, which endpoint served it, or what classification decision routed it there.&lt;/p&gt;

&lt;p&gt;A record that survives an audit looks closer to this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"event_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"01J9Z2Q7K3XN4B8V"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-09-01T09:14:22.481Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"actor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"user_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"u_8812"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"underwriter"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"tenant"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"eu-west"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"classification"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"labels"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"pii"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"financial"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"jurisdiction"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"EU"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"routing"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"policy"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"eu-resident-only"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"endpoint"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"onprem-vllm-01"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"reason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pii_detected"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"llama-3.1-70b-instruct"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"version"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-04-11"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"params"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"temperature"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.2&lt;/span&gt;&lt;span class="p"&gt;}},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"input"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"raw_sha256"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"9f2c..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"redacted"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Applicant [PERSON_1], DTI [NUMBER_1]..."&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"output"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sha256"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"4ab1..."&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"integrity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"prev_hash"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"7d10..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"sig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ed25519:..."&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three details do the heavy lifting. Hashing the raw input lets you prove integrity without storing regulated data twice. Recording the routing decision and its reason is what demonstrates a control was enforced rather than merely documented. Chaining each record to the previous hash is what makes the log tamper-evident instead of just append-only.&lt;/p&gt;

&lt;p&gt;Pin the model version explicitly. "Latest" is not a version, and under the FDA's guidance on &lt;a href="https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-and-machine-learning-software-medical-device" rel="noopener noreferrer"&gt;AI/ML-based Software as a Medical Device&lt;/a&gt;, a clinical model that adapts over time needs a Predetermined Change Control Plan describing what may change and how it gets validated first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where to put it
&lt;/h2&gt;

&lt;p&gt;The enforcement point has to sit between your applications and every model endpoint, because that is the only place where you see all the traffic and can still act before data leaves the perimeter. Classification, redaction, routing and logging happen in one pass, and the routing decision is made before dispatch rather than reconstructed afterwards.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;policies&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;phi"&lt;/span&gt;&lt;span class="pi"&gt;]}&lt;/span&gt;
    &lt;span class="na"&gt;route&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;onprem-clinical&lt;/span&gt;
    &lt;span class="na"&gt;on_missing_endpoint&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;block&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;match&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pii"&lt;/span&gt;&lt;span class="pi"&gt;],&lt;/span&gt; &lt;span class="nv"&gt;jurisdiction&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EU"&lt;/span&gt;&lt;span class="pi"&gt;}&lt;/span&gt;
    &lt;span class="na"&gt;route&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;eu-vpc-endpoint&lt;/span&gt;
    &lt;span class="na"&gt;redact&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;PERSON&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;NATIONAL_ID&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;IBAN&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;default&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;route&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;public-api&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;on_missing_endpoint: block&lt;/code&gt; line matters more than it looks. A policy that silently falls back to a public endpoint when the compliant one is unavailable is not a control. NeuralTrust's writeup on &lt;a href="https://neuraltrust.ai/blog/ai-gateways-data-sovereignty" rel="noopener noreferrer"&gt;how AI gateways enforce data sovereignty&lt;/a&gt; goes deeper on the routing and inspection mechanics, and the &lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;Agent Gateway&lt;/a&gt; product page covers how this is implemented in practice.&lt;/p&gt;

&lt;p&gt;Agents raise the difficulty. A single user request can fan out into many tool calls and model invocations, each a potential egress point, and your audit trail needs to reconstruct the whole chain rather than the first hop. &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;Agent Security&lt;/a&gt; maintains a useful public library of agent threat models and runtime controls if that is where your architecture is heading.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;Three regulatory regimes. One place to enforce them. Build the inference-layer record first, make routing a hard gate rather than a default, and pin versions on everything. The sector-specific paperwork gets much easier once the underlying trail exists.&lt;/p&gt;

&lt;p&gt;For the full regulatory breakdown by sector, the original piece is here: &lt;a href="https://neuraltrust.ai/blog/sovereign-ai-regulated-industries" rel="noopener noreferrer"&gt;Sovereign AI for Regulated Industries&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>agents</category>
      <category>cybersecurity</category>
    </item>
    <item>
      <title>Your Red Team Found a Jailbreak. Now What?</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Tue, 01 Sep 2026 09:58:20 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/your-red-team-found-a-jailbreak-now-what-2god</link>
      <guid>https://dev.to/alessandro_pignati/your-red-team-found-a-jailbreak-now-what-2god</guid>
      <description>&lt;h3&gt;
  
  
  Most AI red teaming tools can break your chatbot. The harder question is what happens to the finding after the scan finishes.
&lt;/h3&gt;

&lt;p&gt;Run any decent adversarial testing tool against an LLM app and you will get findings. Prompt injection lands. The system prompt leaks. A multi-turn escalation gets the agent to call a tool it should never touch. That part of the category is basically solved.&lt;/p&gt;

&lt;p&gt;The part that is not solved is everything after the report renders. Someone reads the PDF, files a ticket, edits a system prompt, and hopes. Three weeks later a model version bumps and nobody knows if the fix held.&lt;/p&gt;

&lt;p&gt;That gap is the actual thing to evaluate when you pick a tool, and it is the lens behind NeuralTrust's &lt;a href="https://neuraltrust.ai/blog/best-ai-red-teaming-platforms" rel="noopener noreferrer"&gt;rundown of the ten platforms worth knowing in 2026&lt;/a&gt;. This is the short version for people who have to wire one of these into a pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why your pentest playbook does not port over
&lt;/h2&gt;

&lt;p&gt;Traditional pentesting assumes a deterministic target. Same input, same code path, same result. You find an injection point, you patch it, it stays patched.&lt;/p&gt;

&lt;p&gt;LLM applications violate all three assumptions. The exploit is natural language, not a malformed packet. The same prompt can pass on Monday and fail on Tuesday because temperature, retrieved context, or a provider-side model update changed underneath you. And the attack surface reopens on every prompt edit, every new tool binding, every RAG index refresh.&lt;/p&gt;

&lt;p&gt;This is not a fringe opinion. OWASP has kept prompt injection at the top of its &lt;a href="https://owasp.org/www-project-top-10-for-large-language-model-applications/" rel="noopener noreferrer"&gt;Top 10 for LLM Applications&lt;/a&gt; for two editions running, and the 2025 document is blunt about why: because of how generative models work, it is unclear whether any foolproof prevention exists. Their recommended mitigation list is defense in depth plus adversarial testing and attack simulation. Not a patch. A control plus continuous verification.&lt;/p&gt;

&lt;p&gt;So a quarterly manual audit is theater. Adversarial testing has to run like a test suite, on every change.&lt;/p&gt;

&lt;h2&gt;
  
  
  The category consolidated fast
&lt;/h2&gt;

&lt;p&gt;Worth knowing before you sign anything, because it changes who owns the roadmap:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Zscaler acquired SPLX in November 2025 and folded it into the Zero Trust Exchange.&lt;/li&gt;
&lt;li&gt;Check Point acquired Lakera in 2025, making it the base of its AI security practice.&lt;/li&gt;
&lt;li&gt;OpenAI announced its acquisition of Promptfoo on March 9, 2026, integrating it into OpenAI Frontier. The open source project stays open under its current license.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is inherently bad. If your SOC already lives in Zscaler or Check Point, consolidation is a feature. But it does mean the red teaming roadmap now follows a much larger platform's priorities. And in the Promptfoo case there is a structural question worth asking out loud: a testing tool owned by a model lab evaluates models in a context shaped by that lab.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things to actually check
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Does a finding become a control?&lt;/strong&gt; This is the one most teams underweight. If your testing vendor and your runtime vendor are different companies, you own the translation layer between them forever. A closed loop looks like: attack lands, finding maps to an enforcement policy, re-run confirms the policy blocks it. NeuralTrust pairs TrustTest with &lt;a href="https://neuraltrust.ai/ai-agent-security" rel="noopener noreferrer"&gt;TrustGuard&lt;/a&gt;, its own inline detection engine, specifically so that handoff is not yours to build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Is it stateful across turns?&lt;/strong&gt; Single-turn scanning is the easy 60 percent. Crescendo, role-play drift, and context hijacking work precisely because each individual message looks fine. If the detection layer evaluates messages in isolation, it will miss the attacks that actually succeed in production. Same question applies to the offensive side: does the tool escalate across a conversation, or just replay a jailbreak corpus?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Does the output survive an audit?&lt;/strong&gt; OWASP, MITRE ATLAS, ISO/IEC 42001, and the EU AI Act are the checkpoints enterprise procurement asks about. A raw JSON dump means someone on your team spends a week translating results into something a risk owner can read.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Does it test agents, not just chat?&lt;/strong&gt; Tool misuse, indirect injection through tool output, memory manipulation, MCP surfaces. The threat model for a tool-calling agent has little to do with the threat model for a chatbot. The teams cataloguing this properly, like the threat library over at &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;agentsecurity.com&lt;/a&gt;, treat agentic failure modes as their own class, and your testing should too.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make it a build step
&lt;/h2&gt;

&lt;p&gt;The practical bar is that a failed attack blocks a merge. Here is roughly what that looks like with TrustTest, which is a Python framework rather than a portal, so it version-controls alongside your app code. Adapted from the &lt;a href="https://docs.neuraltrust.ai/trusttest/getting-started/overview" rel="noopener noreferrer"&gt;official docs&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;trusttest&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;trusttest.catalog.red_team&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;run_red_teaming&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;trusttest.targets.http&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;HttpTarget&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;PayloadConfig&lt;/span&gt;

&lt;span class="n"&gt;target&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;HttpTarget&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://your-api.com/chat&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;payload_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nc"&gt;PayloadConfig&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;message&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;{{ test }}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}),&lt;/span&gt;
    &lt;span class="n"&gt;concatenate_field&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;trusttest&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nb"&gt;type&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;neuraltrust&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;token&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TARGET_TOKEN&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;target_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;TARGET_ID&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;English&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Spanish&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="nf"&gt;run_red_teaming&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;language&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things to notice. The target is an HTTP endpoint, which means this is blackbox. You are testing the deployed application, including its system prompt, retrieval layer, and tool bindings, not an isolated model behind an API. That distinction matters because most real failures live in the glue, not the weights.&lt;/p&gt;

&lt;p&gt;The other is the language loop. Multilingual testing is not a nice-to-have. Safety alignment is unevenly distributed across languages, and a refusal that holds in English frequently does not hold in a lower-resource language. If your app serves non-English users, monolingual testing is measuring the wrong thing. NeuralTrust's own &lt;a href="https://neuraltrust.ai/red-teaming" rel="noopener noreferrer"&gt;red teaming product page&lt;/a&gt; leans on this, and it is one of the more common blind spots in homegrown eval suites.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;Pick on the loop, not the attack count. Every vendor will quote you a big number of adversarial scenarios and every vendor's demo will find something. What separates them is whether a finding turns into an enforced runtime policy and whether a re-run proves the fix held.&lt;/p&gt;

&lt;p&gt;If those are two different vendors, that integration is now your problem. Ask the question during the POC, not after the model version bumps.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>cybersecurity</category>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Your WAF Has No Idea What Your LLM Agent Just Did</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Wed, 26 Aug 2026 14:59:37 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/your-waf-has-no-idea-what-your-llm-agent-just-did-gfh</link>
      <guid>https://dev.to/alessandro_pignati/your-waf-has-no-idea-what-your-llm-agent-just-did-gfh</guid>
      <description>&lt;h3&gt;
  
  
  Why traditional security tooling breaks down for LLM and agent traffic, and what actually needs to sit in the request path
&lt;/h3&gt;

&lt;p&gt;If you've put an AI agent into production, you've probably already noticed the uncomfortable gap. Your WAF is happily passing traffic because every request is a well formed JSON payload. Meanwhile the actual risk, a prompt telling the model to ignore its instructions or call a tool it has no business calling, sails straight through because nothing about it looks malicious at the protocol level.&lt;/p&gt;

&lt;p&gt;That's the core problem. Security tooling built for the API era inspects syntax: headers, payload size, known attack signatures. LLM and agent traffic is an attack surface made of meaning, not syntax, and once you add agents into the mix, it's not even a single request anymore. It's a chain of autonomous decisions, any one of which can be hijacked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why agents make this worse
&lt;/h2&gt;

&lt;p&gt;A plain LLM call that gets manipulated produces bad text. An agent that gets manipulated takes an action. It queries a database, hits an internal API, sends an email. The injected instruction doesn't have to come from the user typing something suspicious either. It can be buried in a PDF the agent retrieves, or hidden in the output of a tool it just called.&lt;/p&gt;

&lt;p&gt;Rate limiting won't catch this because request volume looks completely normal. A compromised agent doing exactly what an attacker wants still fits well within your quota.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where enforcement actually needs to live
&lt;/h2&gt;

&lt;p&gt;The practical answer is a control point that sees every prompt, every tool call, and every completion before it goes anywhere, and applies policy inline rather than logging problems after the fact. That's the job of an AI gateway, and it's a genuinely different job than a guardrails library bolted onto one service, since guardrails have to be reimplemented per application while a gateway enforces consistently across all of them.&lt;/p&gt;

&lt;p&gt;A gateway worth the name should be doing a few concrete things:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Inspecting&lt;/strong&gt; prompts, tool calls, and completions for injection attempts and policy violations in real time&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scoping access&lt;/strong&gt; per token, per agent, per tool, so a single compromised credential doesn't unlock everything&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redacting&lt;/strong&gt; PII in both directions, inbound before it reaches the model, outbound before it reaches the caller&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate limiting&lt;/strong&gt; per user and per agent to contain both abuse and runaway agentic loops&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logging&lt;/strong&gt; everything in a tamper resistant audit trail&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failing closed&lt;/strong&gt; when a check can't complete, rather than letting the request through by default&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last point matters more than it sounds. If your injection classifier times out and the fallback is "allow," you've quietly turned every service degradation into a security hole.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mapping this to a real threat model
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://genai.owasp.org/" rel="noopener noreferrer"&gt;OWASP Top 10 for LLM Applications&lt;/a&gt; is the closest thing the field has to a standard reference here, and it maps cleanly onto gateway controls. Prompt Injection (LLM01) and Sensitive Information Disclosure (LLM02) are the two most obviously suited to inline inspection and redaction. Excessive Agency (LLM06), an agent given more tool access than the task actually requires, is where per-agent, per-tool permissioning earns its keep. Unbounded Consumption (LLM10) is what rate limiting and quotas are for.&lt;/p&gt;

&lt;p&gt;Not everything on that list belongs at the gateway. Supply chain risk and training/fine-tuning data poisoning happen upstream of runtime traffic, so a gateway can restrict which model versions an app is allowed to call, but it's not a substitute for governance further back in the pipeline.&lt;/p&gt;

&lt;h2&gt;
  
  
  What a policy actually looks like
&lt;/h2&gt;

&lt;p&gt;Concretely, this tends to be declarative config rather than code scattered across services:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;policy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;fail_closed&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
  &lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;prompt_injection&lt;/span&gt;
      &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;block&lt;/span&gt;
      &lt;span class="na"&gt;sensitivity&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;high&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pii_detection&lt;/span&gt;
      &lt;span class="na"&gt;scope&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;inbound&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;outbound&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;redact&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;agent_tool_access&lt;/span&gt;
      &lt;span class="na"&gt;allowed_tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;search_docs&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;read_ticket&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;denied_tools&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;send_email&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;execute_code&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
      &lt;span class="na"&gt;require_approval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;update_database&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;rate_limit&lt;/span&gt;
      &lt;span class="na"&gt;scope&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;per_agent&lt;/span&gt;
      &lt;span class="na"&gt;limit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;100&lt;/span&gt;
      &lt;span class="na"&gt;window&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;1m&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The specifics vary by platform, but the shape is always the same: block by default, scope tools tightly, redact before data leaves your boundary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compliance is starting to demand this anyway
&lt;/h2&gt;

&lt;p&gt;This isn't purely a "nice to have" argument anymore. The EU AI Act requires high risk AI systems to support automatic event logging (Article 12) and requires providers and deployers to retain those logs, which pushes audit trails from best practice into a compliance requirement for a lot of teams shipping agentic systems into the EU.&lt;/p&gt;

&lt;h2&gt;
  
  
  Enforcement vs. assessment
&lt;/h2&gt;

&lt;p&gt;It's worth separating two things that get conflated a lot: enforcing policy inline versus assessing agent risk and running adversarial tests against your workflows to figure out what that policy should even be. NeuralTrust splits these into &lt;a href="https://neuraltrust.ai/ai-gateway" rel="noopener noreferrer"&gt;TrustGate&lt;/a&gt;, the inline enforcement gateway, and TrustGuard, the governance and threat detection layer that informs it. If you want the fuller picture on agent-specific threats and how gateways fit into securing them, NeuralTrust's &lt;a href="https://neuraltrust.ai/blog/how-to-secure-llm-and-agent-traffic" rel="noopener noreferrer"&gt;original writeup on this topic&lt;/a&gt; and their &lt;a href="https://neuraltrust.ai/ai-agent-security" rel="noopener noreferrer"&gt;AI agent security overview&lt;/a&gt; are both worth a look, as is &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;agentsecurity.com&lt;/a&gt; if you're mapping out the broader landscape of tooling in this space.&lt;/p&gt;

&lt;p&gt;The bottom line: if your security stack can't read the meaning of a prompt or the intent behind a tool call, it isn't actually securing your LLM and agent traffic, it's just watching it go by.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>agents</category>
    </item>
    <item>
      <title>Your AI Gateway Isn't Watching Your Agent's Tool Calls. Here's Why That Matters.</title>
      <dc:creator>Alessandro Pignati</dc:creator>
      <pubDate>Wed, 26 Aug 2026 08:59:41 +0000</pubDate>
      <link>https://dev.to/alessandro_pignati/your-ai-gateway-isnt-watching-your-agents-tool-calls-heres-why-that-matters-kh8</link>
      <guid>https://dev.to/alessandro_pignati/your-ai-gateway-isnt-watching-your-agents-tool-calls-heres-why-that-matters-kh8</guid>
      <description>&lt;h3&gt;
  
  
  A practical breakdown of what an AI gateway actually sees versus what an MCP gateway sees, and why production agents usually need both
&lt;/h3&gt;

&lt;p&gt;Picture a fairly normal setup. You've got an agent behind an AI gateway. Prompts are rate limited, PII gets scrubbed, costs are tracked per team, jailbreak attempts get flagged. Solid setup. Then the agent picks up tool calling through MCP so it can hit your CRM, your internal search index, a couple of internal APIs. Nothing changes on the gateway side because, as far as the gateway is concerned, nothing changed. The model call still looks the same going in and coming out.&lt;/p&gt;

&lt;p&gt;Except now there's a whole second leg of the request that nobody is watching.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two different jobs, one request path
&lt;/h2&gt;

&lt;p&gt;An AI gateway is a reverse proxy for LLM traffic. It sits between your app and whichever model provider you're calling (OpenAI, Anthropic, Bedrock, Gemini, self-hosted, doesn't matter) and normalizes all of that behind one interface. The stuff it's built for is specific to model traffic: token-based rate limits instead of raw request counts, streaming responses, provider fallback when one goes down, semantic caching, cost attribution, and content policy on the prompt going in and the completion coming out.&lt;/p&gt;

&lt;p&gt;What it does not do, structurally, is see what happens after the model decides to call a tool. Once a completion includes a tool call, execution moves to whatever runs your agent's tool logic, and if that's happening over the Model Context Protocol, it's a different transport entirely.&lt;/p&gt;

&lt;p&gt;That's the gap an MCP gateway is built to close. MCP is an open standard for letting an agent discover and call external tools and data sources through a consistent client-server interface, instead of every team hand-rolling a connector for every model-to-system combination. It's often described as turning an M×N integration problem (M agents, N systems, custom glue for every pair) into an M+N one, one client, one server per system, and anything MCP-compliant can talk to anything else MCP-compliant.&lt;/p&gt;

&lt;p&gt;Convenient, but MCP itself has no opinion on governance. Nothing in the protocol stops an agent from calling every tool a server exposes, and nothing logs what it actually did with them. An MCP gateway sits in front of that traffic and adds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Auth for every agent-to-server connection, OAuth flows for remote servers included&lt;/li&gt;
&lt;li&gt;Per-tool, per-agent authorization, not just "can this agent reach the server" but "which specific tools on it"&lt;/li&gt;
&lt;li&gt;Inventory of MCP servers across the org, including the shadow ones nobody registered&lt;/li&gt;
&lt;li&gt;Audit logs of who called what, with what arguments, and what came back&lt;/li&gt;
&lt;li&gt;Inspection of tool calls and results for injection attempts or exfiltration hiding in a response&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What that actually looks like
&lt;/h2&gt;

&lt;p&gt;A rough shape for a per-tool policy an MCP gateway would enforce:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;agent&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;support-bot&lt;/span&gt;
&lt;span class="na"&gt;server&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;internal-crm&lt;/span&gt;
&lt;span class="na"&gt;allow&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;tool&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;lookup_customer&lt;/span&gt;
    &lt;span class="na"&gt;max_calls_per_minute&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;20&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;tool&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;update_ticket_status&lt;/span&gt;
    &lt;span class="na"&gt;require_approval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
&lt;span class="na"&gt;deny&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;tool&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;export_customer_data&lt;/span&gt;
&lt;span class="na"&gt;audit&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;all&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;None of that is something an AI gateway has any hooks for, because it's not model traffic. It's a different policy surface with a different unit of enforcement (per tool, per server, per agent identity) versus the AI gateway's unit (per model, per provider, per token).&lt;/p&gt;

&lt;h2&gt;
  
  
  So which one do you need
&lt;/h2&gt;

&lt;p&gt;If your agent only sends prompts to a model and never touches a tool, an AI gateway on its own probably covers you. The moment it calls into MCP servers to query a database, hit an API, or take an action on an external system, you've got a blind spot without MCP-layer controls. Most production agents that do real work fall into the second bucket, which is why most serious deployments end up running both layers, ideally on shared infrastructure so policy isn't defined twice and drifting apart.&lt;/p&gt;

&lt;p&gt;This is roughly the bet behind &lt;a href="https://github.com/NeuralTrust/TrustGate" rel="noopener noreferrer"&gt;NeuralTrust's TrustGate&lt;/a&gt;, an open-source gateway handling LLM, MCP, and agent-to-agent traffic under one control plane. The reasoning holds regardless of which vendor or open-source project you land on: a threat that gets past the model-traffic layer can still get caught at the tool-traffic layer, and vice versa, but only if something is actually watching both. It's the same instinct behind the broader move toward dedicated &lt;a href="https://agentsecurity.com/" rel="noopener noreferrer"&gt;agent security&lt;/a&gt; tooling rather than trying to stretch a general API gateway or a model-only gateway to cover agentic behavior it wasn't designed to see.&lt;/p&gt;

&lt;h2&gt;
  
  
  The standards bodies are catching up
&lt;/h2&gt;

&lt;p&gt;This isn't just a vendor talking point anymore. In December 2025, Anthropic donated MCP to the newly formed Agentic AI Foundation, a directed fund under the Linux Foundation co-founded with Block and OpenAI, specifically to keep governance of the protocol vendor-neutral as adoption scales. And in February 2026, &lt;a href="https://www.nist.gov/node/1906621" rel="noopener noreferrer"&gt;NIST's Center for AI Standards and Innovation launched an AI Agent Standards Initiative&lt;/a&gt;, with one of its three pillars aimed squarely at agent identity, authentication, and authorization. Two separate signals pointing at the same conclusion. The controls an MCP gateway provides today (identity, per-tool auth, audit trails) are on a trajectory to becoming baseline expectations, not optional extras.&lt;/p&gt;

&lt;p&gt;If you want the deeper dive on AI gateways specifically, including where they overlap with plain API gateways, &lt;a href="https://neuraltrust.ai/blog/what-is-an-ai-gateway" rel="noopener noreferrer"&gt;NeuralTrust has a longer breakdown here&lt;/a&gt;. And the &lt;a href="https://neuraltrust.ai/blog/ai-gateway-vs-mcp-gateway" rel="noopener noreferrer"&gt;original comparison this post is based on&lt;/a&gt; goes further into the architecture diagrams and a full feature-by-feature breakdown if you want the long version.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Bottom line:&lt;/strong&gt; an AI gateway secures the thinking. An MCP gateway secures the acting. If your agent does both, so should your infrastructure.&lt;/p&gt;

</description>
      <category>machinelearning</category>
      <category>ai</category>
      <category>cybersecurity</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
