<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dinesh Kumar Sarangapani</title>
    <description>The latest articles on DEV Community by Dinesh Kumar Sarangapani (@dks).</description>
    <link>https://dev.to/dks</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3113193%2F7ba33a44-54e6-44ac-9be0-79f66d6f437b.jpg</url>
      <title>DEV Community: Dinesh Kumar Sarangapani</title>
      <link>https://dev.to/dks</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dks"/>
    <language>en</language>
    <item>
      <title>Building Scalable Agentic AI Platforms: A Technical Deep Dive - Part 2</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 16:07:03 +0000</pubDate>
      <link>https://dev.to/dks/building-scalable-agentic-ai-platforms-a-technical-deep-dive-part-2-5958</link>
      <guid>https://dev.to/dks/building-scalable-agentic-ai-platforms-a-technical-deep-dive-part-2-5958</guid>
      <description>&lt;p&gt;Before diving into Part 2, make sure to read &lt;a href="https://dev.to/blog/2025/03/building-scalable-ai-platforms/"&gt;Part 1&lt;/a&gt; where we covered the fundamentals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Module 1: Model Management using LLM Gateway
&lt;/h2&gt;

&lt;p&gt;There are several models that are useful for specific purposes. Even though there are general-purpose models, we need to choose the best-performing model for our task. Deploying those models in various cloud providers and applying the security principles, tracing, cost controlling them is a Challenge. &lt;/p&gt;

&lt;p&gt;&lt;a href="/images/agentic-ai/before_llm-gateway.png" class="article-body-image-wrapper"&gt;&lt;img src="/images/agentic-ai/before_llm-gateway.png" alt="LLM Gateway Architecture Before"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After the gateway,&lt;/p&gt;

&lt;p&gt;&lt;a href="/images/agentic-ai/after-llm-gw.png" class="article-body-image-wrapper"&gt;&lt;img src="/images/agentic-ai/after-llm-gw.png" alt="LLM Gateway Architecture After"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Key features:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Routing to right Model&lt;/li&gt;
&lt;li&gt;Central logging for Compliance checks&lt;/li&gt;
&lt;li&gt;Metrics collection for cost and usage&lt;/li&gt;
&lt;/ol&gt;




</description>
      <category>ai</category>
      <category>architecture</category>
      <category>llm</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Building an Enterprise MCP Tool Broker: Turning OpenAPI into Agent Tools</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:24:32 +0000</pubDate>
      <link>https://dev.to/dks/building-an-enterprise-mcp-tool-broker-turning-openapi-into-agent-tools-22og</link>
      <guid>https://dev.to/dks/building-an-enterprise-mcp-tool-broker-turning-openapi-into-agent-tools-22og</guid>
      <description>&lt;p&gt;AI agents are only as useful as the tools they can call. An agent without tools can only chat. An agent with tools can query databases, create tickets, inspect files, and trigger deployments.&lt;/p&gt;

&lt;p&gt;In most enterprises, those capabilities already exist as REST APIs. Over the last decade, organizations built hundreds of internal microservices documented with OpenAPI.&lt;/p&gt;

&lt;p&gt;When we started connecting agents to our internal systems, we faced a choice. Should we manually write dedicated tool wrappers for every API? Or could we build a broker that translates existing OpenAPI specifications into agent tools automatically?&lt;/p&gt;

&lt;p&gt;We chose to build a centralized tool broker using the Model Context Protocol (MCP). Here is how the idea works, how it performed in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: A Dynamic Tool Broker
&lt;/h2&gt;

&lt;p&gt;Writing dedicated tool servers for hundreds of microservices does not scale. Each engineering team would need to learn a new protocol, maintain extra code, and keep tool schemas manually synchronized with API updates.&lt;/p&gt;

&lt;p&gt;The core idea was to build a single enterprise tool broker.&lt;/p&gt;

&lt;p&gt;The broker sits between AI agents and internal APIs. To the agent, the broker looks like an MCP server offering a unified catalog of tools. Behind the scenes, the broker ingests existing OpenAPI specifications from registered services and dynamically translates each API operation into a typed, callable tool.&lt;/p&gt;

&lt;p&gt;When an agent invokes a tool, the broker validates the caller's identity, translates the arguments into an HTTP request, calls the backend service, and returns the formatted result back to the model.&lt;/p&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Zero Work for Service Teams&lt;/strong&gt;: Existing microservices did not need code changes. As long as a team maintained an accurate OpenAPI specification, their service became agent-callable instantly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standardized Tool Discovery&lt;/strong&gt;: Instead of teaching every agent about individual services, agents connect to one broker endpoint. The broker dynamically filters which tools an agent can see based on the caller's permissions and role.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decoupled Architecture&lt;/strong&gt;: Upstream microservices remain pure REST APIs. If the agent protocol changes or new client interfaces emerge, the translation logic lives in the broker rather than scattered across dozens of repositories.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governed Access Controls&lt;/strong&gt;: The broker enforces centralized authorization before executing any backend request. We can enforce least-privilege access, audit tool invocation logs, and disable high-risk actions without modifying the underlying microservices.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Vague Endpoint Summaries Confuse Models&lt;/strong&gt;: Large language models pick tools based on natural language descriptions. If an API has a generic description like &lt;code&gt;Process data&lt;/code&gt;, the model will either hallucinate its purpose or never call it. Clear, descriptive summaries in your API specs are critical for agent accuracy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schema Reference Inlining (&lt;code&gt;$ref&lt;/code&gt;)&lt;/strong&gt;: Many enterprise API specifications use internal schema references (&lt;code&gt;$ref&lt;/code&gt;) to share data types across endpoints. Many LLM engines and agent frameworks cannot resolve external reference pointers automatically. The broker must resolve and inline referenced schemas before handing tool definitions to the model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Circular Schemas&lt;/strong&gt;: Real-world enterprise APIs frequently contain circular schema definitions (for example, a &lt;code&gt;Folder&lt;/code&gt; that contains a list of &lt;code&gt;Folder&lt;/code&gt; items). If your schema resolver does not set recursion depth limits, it will enter an infinite loop or produce massive schemas that exhaust the model's context window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context Window Blowups from Large Payloads&lt;/strong&gt;: A GET endpoint that returns 10 megabytes of JSON is fine for a frontend data table, but it will overflow an LLM's context window. The broker must detect oversized responses, summarize them, or paginate results before returning them to the agent.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write Actions Need Confirmation&lt;/strong&gt;: Read-only tools (searching data, fetching records) are safe to execute autonomously. Write tools (deleting resources, triggering updates) should require explicit human confirmation or stricter permission gates before execution.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>api</category>
      <category>architecture</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Responsible AI for Autonomous Systems: Principles Beyond Traditional Chatbots</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:22:42 +0000</pubDate>
      <link>https://dev.to/dks/responsible-ai-for-autonomous-systems-principles-beyond-traditional-chatbots-2ekf</link>
      <guid>https://dev.to/dks/responsible-ai-for-autonomous-systems-principles-beyond-traditional-chatbots-2ekf</guid>
      <description>&lt;p&gt;Most Responsible AI (RAI) frameworks were built for an earlier generation of machine learning.&lt;/p&gt;

&lt;p&gt;They focused on predictive models (like credit scoring or fraud detection) or passive chatbots. Frameworks from NIST, Microsoft, Google, and Anthropic established valuable principles: fairness, reliability, privacy, transparency, and accountability.&lt;/p&gt;

&lt;p&gt;Those principles remain essential. But when you move from passive chatbots to autonomous, tool-using AI agents, traditional principles are no longer enough.&lt;/p&gt;

&lt;p&gt;A standard chatbot only generates text. If it hallucinates an inaccurate sentence, the user can read it and spot the error.&lt;/p&gt;

&lt;p&gt;An AI agent does not just output text. It plans steps, selects tools, writes queries, triggers API calls, and updates external systems. If an agent hallucinates a step or falls victim to prompt injection, it can delete customer data, exhaust cloud budgets, or trigger cascading system failures.&lt;/p&gt;

&lt;p&gt;To build an enterprise agentic platform safely, we expanded traditional responsible AI tenets into specific principles for autonomous systems.&lt;/p&gt;

&lt;p&gt;Here is the idea, how it worked in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: Bounding Autonomous Action
&lt;/h2&gt;

&lt;p&gt;An autonomous agent must never operate without hard boundaries. We structured our agentic responsible AI principles across three core categories:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+--------------------------------------------------------------------------+
|                  Responsible AI for Autonomous Systems                   |
|                                                                          |
|  1. Bounding &amp;amp; Control    ──▶ Restrict what an agent is allowed to do.   |
|  2. Reliability &amp;amp; Trust   ──▶ Ensure reasoning is grounded and clear.    |
|  3. Security &amp;amp; Alignment  ──▶ Protect against attacks and resource run.  |
+--------------------------------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1. Bounded Agency
&lt;/h3&gt;

&lt;p&gt;An agent's autonomy must be strictly limited to defined, verifiable boundaries. Its planning and tool use must never expand beyond its intended purpose.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Enforce boundaries through code and architecture, not just by writing gentle instructions in a system prompt.&lt;/li&gt;
&lt;li&gt;Restrict tool access using the principle of least privilege. If an agent only needs to read records, it must not have access to an API endpoint that can modify or delete records.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Economic Boundedness
&lt;/h3&gt;

&lt;p&gt;Agents must operate within hard computational and financial budgets.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Without economic limits, an agent caught in a recursive planning loop can spawn thousands of API calls, leading to a "Denial of Wallet" attack or massive cloud bills.&lt;/li&gt;
&lt;li&gt;Every agent session must enforce strict caps on maximum reasoning turns, token consumption per request, and total financial cost per day.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Predictable Tool Use
&lt;/h3&gt;

&lt;p&gt;Agents must only invoke pre-approved tools with strictly validated input parameters.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tool arguments generated by a language model should be treated like untrusted user input.&lt;/li&gt;
&lt;li&gt;A deterministic mediation proxy must validate the arguments against strict JSON schemas before forwarding the call to any backend service.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4. Human-in-the-Loop Oversight &amp;amp; Reversibility
&lt;/h3&gt;

&lt;p&gt;Critical and destructive actions must require human confirmation.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If an agent suggests an action with high impact (such as deleting a resource, modifying financial records, or sending an external communication), the platform halts execution and presents an explicit approval card to a human operator.&lt;/li&gt;
&lt;li&gt;Tasks must have clear termination criteria ("definition of done"). If an action can be undone, provide an automated rollback mechanism.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. Verifiable Groundedness &amp;amp; Operational Explainability
&lt;/h3&gt;

&lt;p&gt;An agent must never take action based on fabricated information.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;All factual claims and decision-making context must be grounded in verified enterprise documents retrieved via secure RAG pipelines.&lt;/li&gt;
&lt;li&gt;The agent must be able to "show its work" on demand, articulating not just what action it chose, but why it selected that step over alternatives.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Preventing Runaway Reasoning Loops&lt;/strong&gt;: Enforcing economic boundedness saved thousands of dollars in cloud spend. When complex edge cases caused agents to enter recursive planning cycles, hard token and turn limits stopped the process gracefully rather than exhausting budgets.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Defending Against Excessive Agency&lt;/strong&gt;: Applying least-privilege tool allowlists stopped prompt injection attacks from causing real-world damage. Even when an untrusted document tricked an agent into attempting a system change, the platform blocked the call because the tool was not in the agent's authorized allowlist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Building User and Operator Confidence&lt;/strong&gt;: Having human-in-the-loop checkpoints for write actions made business teams willing to adopt agents. Users knew the agent could not execute destructive actions without explicit, manual sign-off.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fast Incident Auditing&lt;/strong&gt;: When an agent produced an unexpected output, operational explainability allowed engineers to inspect the exact reasoning chain, tool parameters, and retrieved document chunks in minutes.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Human Review Fatigue&lt;/strong&gt;: If you require human approval for every minor, trivial action, users will stop reading the details and blindly click "Approve." Only require human-in-the-loop gates for truly consequential, irreversible, or write-enabled operations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Agent Coordination Deadlocks&lt;/strong&gt;: When multiple autonomous agents interact (such as an orchestrator delegating tasks to subagents), they can enter circular dependency loops. Implement timeout limits and centralized supervisors to deconflict competing actions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Relying on Prompt Constraints Alone&lt;/strong&gt;: Never trust a system prompt like &lt;em&gt;"Do not delete records"&lt;/em&gt; as your sole security boundary. Models can be manipulated by clever jailbreaks or indirect prompt injections. Hard constraints must live in deterministic backend code and API permissions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stale World Models&lt;/strong&gt;: An agent that relies on cached information might attempt to use an API that is offline or reference an item that has already been moved. Ensure agents verify state before executing critical actions.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Why Traditional API Gateways Break with Generative AI</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:22:39 +0000</pubDate>
      <link>https://dev.to/dks/why-traditional-api-gateways-break-with-generative-ai-2jj7</link>
      <guid>https://dev.to/dks/why-traditional-api-gateways-break-with-generative-ai-2jj7</guid>
      <description>&lt;p&gt;For years, enterprise architecture followed a standard rule: place an enterprise API management gateway in front of every backend service.&lt;/p&gt;

&lt;p&gt;The pattern was proven for REST and GraphQL microservices. A client made a request, the gateway verified a token, applied a rate limit like 100 requests per minute, forwarded the call to a database, and returned an atomic JSON response in 50 milliseconds.&lt;/p&gt;

&lt;p&gt;When organizations place those same legacy gateways in front of generative AI models and autonomous agents, the architecture breaks down.&lt;/p&gt;

&lt;p&gt;This post explains why traditional API gateways struggle with generative AI workloads, how a streaming-native edge layer solves the problem, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem: The Generative AI Mismatch
&lt;/h2&gt;

&lt;p&gt;Traditional API gateways were designed around three core assumptions that are invalid for LLMs:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Atomic Responses vs. Real-Time Token Streams&lt;/strong&gt;: Traditional proxies often buffer the entire HTTP response in memory to inspect headers, check data leakage, or calculate compression. LLMs generate text token by token. Buffering destroys the real-time typing experience, causes buffer overflows on long responses, and triggers premature gateway timeouts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Short-Lived Connections vs. Minute-Long Sessions&lt;/strong&gt;: Traditional microservice requests complete in milliseconds. AI agent reasoning, multi-step tool execution, and complex synthesis can hold an HTTP connection open for 30 to 120 seconds. In high-concurrency environments, long-lived connections quickly exhaust gateway connection pools, causing false timeouts even when server CPU utilization is low.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Request Counting vs. Token Consumption&lt;/strong&gt;: Standard gateways rate-limit by request count (e.g., 60 requests per minute). But in generative AI, one request might be a 10-token greeting, while another is an 80,000-token document analysis costing 1,000 times more compute. Request-based rate limiting fails to protect upstream quotas or prevent cost overruns.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  The Idea: A Streaming-Native Edge Architecture
&lt;/h2&gt;

&lt;p&gt;Instead of forcing generative AI traffic through a legacy API gateway, the idea was to separate traditional API management from streaming AI traffic.&lt;/p&gt;

&lt;p&gt;We built a lightweight, streaming-first edge proxy tailored to the mechanics of AI protocols:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Zero-Buffering Passthrough&lt;/strong&gt;: The edge proxy authenticates the caller during the initial HTTP handshake and immediately steps out of the data path, streaming raw bytes directly from the upstream model service to the client.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token and Concurrency Quotas&lt;/strong&gt;: Rather than counting raw requests, the rate limiter tracks active concurrent runs and cumulative token usage, throttling tenants when they approach usage thresholds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heartbeat Management&lt;/strong&gt;: For reasoning models that think for several seconds before emitting their first token, the edge layer emits periodic keep-alive comment frames to prevent intermediate corporate proxies and firewalls from dropping the idle TCP connection.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Instant User Responsiveness&lt;/strong&gt;: Eliminating gateway response buffering reduced latency to the first visible token and restored a smooth streaming typing effect for end users.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Resilient Connection Pools&lt;/strong&gt;: By handling connections asynchronously with lightweight event loops instead of thread-per-connection pools, the edge layer supported thousands of concurrent long-lived streams without degrading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fair Usage and Cost Protection&lt;/strong&gt;: Token-based rate limiting prevented runaway scripts or high-volume batch queries from starving interactive users or exhausting cloud provider quotas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clean Operational Separation&lt;/strong&gt;: Platform engineers retained full control over AI-specific edge policies without being constrained by enterprise-wide API gateway release cycles.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Corporate Intermediary Proxies&lt;/strong&gt;: Even if your edge proxy streams without buffering, intermediate corporate proxies, VPNs, or Web Application Firewalls (WAFs) in the user's network might still buffer responses. Explicitly disable buffering headers (&lt;code&gt;X-Accel-Buffering: no&lt;/code&gt;) on every streaming response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent Client Disconnects&lt;/strong&gt;: When a user closes their browser tab mid-stream, the connection drops. If your edge layer does not actively listen for client disconnects, your backend model service will keep generating tokens and burning money for an audience that is no longer there. Ensure client disconnects immediately propagate upstream to cancel the inference run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token Accounting Lag&lt;/strong&gt;: Accurately counting tokens requires reading the final usage metadata at the very end of the stream. If you only enforce quotas synchronously before a request starts, a user can initiate several massive concurrent requests that exceed their allotment before the first one finishes. Combine pre-request concurrency limits with post-request token accounting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Heartbeat Hygiene&lt;/strong&gt;: Keep-alive frames are vital for long-thinking models, but your client-side parsers must know how to ignore comment frames (&lt;code&gt;: keepalive\n\n&lt;/code&gt;) without rendering them as garbled text in the UI.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Resilient Multi-Cloud LLM Gateways: Routing, Rate Limits, and Failovers</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:19:26 +0000</pubDate>
      <link>https://dev.to/dks/resilient-multi-cloud-llm-gateways-routing-rate-limits-and-failovers-1pe</link>
      <guid>https://dev.to/dks/resilient-multi-cloud-llm-gateways-routing-rate-limits-and-failovers-1pe</guid>
      <description>&lt;p&gt;If your enterprise AI application connects directly to a single cloud provider's LLM endpoint, your application will eventually experience downtime.&lt;/p&gt;

&lt;p&gt;All foundation model APIs suffer from operational challenges:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Global capacity shortages during peak hours.&lt;/li&gt;
&lt;li&gt;Sudden rate-limit errors when token bucket quotas are exhausted.&lt;/li&gt;
&lt;li&gt;Regional latency spikes where Time To First Token jumps from 400 milliseconds to several seconds.&lt;/li&gt;
&lt;li&gt;Partial brownouts where streaming connections hang indefinitely without emitting bytes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When critical internal systems and production users rely on AI agents, a cloud provider outage cannot be allowed to halt business operations.&lt;/p&gt;

&lt;p&gt;To solve this, we architected a resilient multi-cloud model gateway. It presents a standard OpenAI-compatible interface to internal applications while dynamically orchestrating requests across multiple cloud providers.&lt;/p&gt;

&lt;p&gt;Here is the idea, how it performed in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: A Resilient Model Proxy
&lt;/h2&gt;

&lt;p&gt;Rather than letting each internal application talk directly to external LLM vendors, all traffic routes through a centralized gateway service.&lt;/p&gt;

&lt;p&gt;The gateway manages three core responsibilities:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Dynamic Provider Routing&lt;/strong&gt;: Applications request a logical model capability (such as general-purpose completion or complex reasoning). The gateway maps that capability to healthy providers and regions based on priority, latency, and cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Circuit Breakers and Fast Failover&lt;/strong&gt;: If a primary cloud provider returns errors or times out, the gateway immediately fails over to an alternative provider or region before the user experiences an error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upstream Rate Limit Coordination&lt;/strong&gt;: The gateway tracks provider capacity and respects throttling signals so that traffic spikes do not cause cascading failures.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;High Availability via Automated Failover&lt;/strong&gt;: When one cloud region experienced degraded capacity, the gateway automatically diverted subsequent requests to healthy regions or secondary cloud providers. End users experienced a brief delay of a few hundred milliseconds rather than a failed run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standardized Developer Experience&lt;/strong&gt;: Application teams developed against one consistent API contract. They did not need to write custom SDK code for each cloud provider, and changing underlying models required zero application code modifications.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Throttling Protection with Rate-Limit Forwarding&lt;/strong&gt;: When an upstream provider returned an HTTP 429 response, the gateway captured the upstream wait window and forwarded it to clients while temporarily tripping a local circuit breaker. This prevented retry storms from overwhelming overloaded providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Actionable Latency Observability&lt;/strong&gt;: Instead of tracking total request time, the gateway isolated &lt;strong&gt;Time To First Token (TTFT)&lt;/strong&gt; from &lt;strong&gt;Time Per Output Token (TPOT)&lt;/strong&gt;. This allowed on-call engineers to distinguish between upstream scheduling queues and slow token generation throughput.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Prompt Replay on Streaming Failures&lt;/strong&gt;: If an upstream provider fails &lt;em&gt;before&lt;/em&gt; emitting any tokens, the gateway can cleanly replay the prompt to a secondary provider. However, if the stream fails &lt;em&gt;mid-generation&lt;/em&gt; after emitting 50 tokens, replaying the entire prompt from scratch will duplicate text in the user's chat. The gateway must detect whether bytes have traversed to the client and handle partial stream interruptions gracefully.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Behavioral Divergence Across Models&lt;/strong&gt;: Even when two models are considered comparable, their system prompt interpretation, tool calling formats, and output schemas may subtly differ. Ensure your agents are tested against the fallback model families so that an automated failover does not produce unexpected tool argument syntax.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Streaming Deadlocks and Connection Leaks&lt;/strong&gt;: When users abandon a web page or cancel a request mid-stream, client connections close. If your gateway does not actively propagate cancellation signals upstream, the underlying cloud connection remains open, continuing to consume model tokens and holding sockets in your connection pool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Circuit Breaker Flapping&lt;/strong&gt;: If your circuit breaker threshold is too aggressive, a few transient network drops will unnecessarily route all traffic to more expensive backup providers. Use a combination of consecutive failure counts and cooldown intervals with canary probe requests before fully reopening circuits.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>api</category>
      <category>architecture</category>
      <category>cloud</category>
      <category>llm</category>
    </item>
    <item>
      <title>Network Segmentation for AI Systems: Containing Blast Radiuses in Agent Architectures</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:19:23 +0000</pubDate>
      <link>https://dev.to/dks/network-segmentation-for-ai-systems-containing-blast-radiuses-in-agent-architectures-3lkh</link>
      <guid>https://dev.to/dks/network-segmentation-for-ai-systems-containing-blast-radiuses-in-agent-architectures-3lkh</guid>
      <description>&lt;p&gt;When teams build an initial AI system, they usually launch all their components into a single virtual network.&lt;/p&gt;

&lt;p&gt;The orchestrator, vector databases, cache clusters, and background workers share the same private subnets. Everything communicates over internal IP addresses without friction. Routing is simple, and there is no inter-network peering to manage.&lt;/p&gt;

&lt;p&gt;For traditional internal apps, this flat topology is common. For agentic AI systems, it creates a serious security risk.&lt;/p&gt;

&lt;p&gt;Autonomous AI agents do something standard microservices rarely do: they dynamically construct network requests, invoke external tools, and fetch content from untrusted endpoints based on model reasoning.&lt;/p&gt;

&lt;p&gt;If an agent runs in a flat network and encounters prompt injection or malicious input, an attacker can use the agent as an internal proxy to explore your private infrastructure.&lt;/p&gt;

&lt;p&gt;This post explains the network threat model for AI agents, how a zoned network design contains blast radiuses, how it works in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Problem: The Agentic Threat Model
&lt;/h2&gt;

&lt;p&gt;In traditional software, network requests are deterministic. Code specifies exact hosts, ports, and payload structures.&lt;/p&gt;

&lt;p&gt;AI agents are non-deterministic. A user prompt or a retrieved document can instruct an agent to call an HTTP tool with arbitrary parameters.&lt;/p&gt;

&lt;p&gt;This opens the door to &lt;strong&gt;Server-Side Request Forgery (SSRF)&lt;/strong&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;An agent retrieves an external document or processes user input containing hidden instructions.&lt;/li&gt;
&lt;li&gt;The instructions direct the agent to fetch an internal URL (such as a cloud instance metadata service or an internal database management port).&lt;/li&gt;
&lt;li&gt;The agent executes the request using its host identity.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;In a flat network, the agent shares private network routes with critical databases, internal credentials, and other services. Once an agent is compromised, the attacker inherits the agent's internal network access.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: Zoned Network Micro-Perimeters
&lt;/h2&gt;

&lt;p&gt;To protect against this, production AI platforms should adopt a zoned network architecture. Instead of placing everything together, components are grouped by their trust level and access requirements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The Execution Zone&lt;/strong&gt;: Hosts the core agent orchestration engine and ephemeral session state. It can query vector indexes and send prompts to the model gateway, but has no direct internet egress and no route to raw backend infrastructure.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Protected Zone&lt;/strong&gt;: Houses sensitive vector embeddings, semantic indexes, and reference datasets. Inbound access is strictly restricted to authorized queries from the execution zone. It has zero public internet access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Sandboxed Zone&lt;/strong&gt;: Dedicated to parsing untrusted external files (PDFs, spreadsheets, HTML). Document parsing is historically prone to memory-safety bugs and zero-day parser vulnerabilities. By running parsers in an isolated network enclave, untrusted parser code never runs in the same network space as operational agents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The Egress Zone&lt;/strong&gt;: All external communication (calling model APIs or executing approved web tools) flows through an inspected egress layer with stateful firewalls.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Strict Blast Radius Containment&lt;/strong&gt;: If a prompt injection tricks an agent into attempting an unauthorized network scan, the network route simply does not exist. The packet is dropped at the virtual network boundary before it can reach internal databases.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;True Defense-in-Depth&lt;/strong&gt;: Application-level input validation (such as URL regexes) can be bypassed by DNS rebinding, alternative IP encodings, or HTTP redirect chains. Enforcing drop rules at the network firewall guarantees that traffic to private IP spaces and metadata services is blocked regardless of what the application code does.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Isolated File Ingestion&lt;/strong&gt;: Separating document parsing into a sandboxed network prevented vulnerabilities in third-party document processing libraries from compromising the core agent runtime or accessing user session stores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Governed Managed Services&lt;/strong&gt;: Using private interface endpoints (such as PrivateLink) kept database and storage traffic entirely on the cloud provider's private physical backbone, eliminating exposure to public internet routing.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;DNS Rebinding Attacks&lt;/strong&gt;: An attacker can provide a domain name that resolves to a public IP during initial application checks, but points to a private internal IP by the time the agent's HTTP client connects. Ensure your egress layer resolves DNS and validates the destination IP at connection time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hidden HTTP Redirect Chains&lt;/strong&gt;: A public URL may return an HTTP 302 redirect pointing to an internal IP or cloud metadata service. Ensure your HTTP client libraries or egress firewalls inspect redirect targets before following them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cross-Zone Network Costs&lt;/strong&gt;: Moving traffic across peered virtual networks or transit gateways incurs cloud data transfer fees. High-volume document embedding pipelines can generate terabytes of data. Keep tightly coupled services in the same region and optimize payload transfers to avoid unexpected network bills.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Operational Overhead of Peering&lt;/strong&gt;: Managing multiple isolated virtual networks with security groups and route tables manually is unsustainable. Define all networking zones in Infrastructure as Code (IaC) so that environments remain reproducible and auditable across development, staging, and production.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>networking</category>
      <category>security</category>
    </item>
    <item>
      <title>The Responsible AI Lifecycle: Select, Test, Integrate, and Operate</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:17:21 +0000</pubDate>
      <link>https://dev.to/dks/the-responsible-ai-lifecycle-select-test-integrate-and-operate-5bfe</link>
      <guid>https://dev.to/dks/the-responsible-ai-lifecycle-select-test-integrate-and-operate-5bfe</guid>
      <description>&lt;p&gt;When enterprise technology providers like Microsoft, Google, Meta, and Anthropic publish Responsible AI guidelines, they offer powerful safety tools: content filters, prompt shields, and toxicity classifiers.&lt;/p&gt;

&lt;p&gt;However, relying solely on provider tools leaves a critical &lt;strong&gt;responsibility gap&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Most vendor-supplied safety tools act as &lt;strong&gt;reactive filters&lt;/strong&gt;. They catch harmful text or known jailbreak patterns during live execution.&lt;/p&gt;

&lt;p&gt;They do not tell you if a candidate foundation model is prone to subtle bias in your specific business domain. They do not test whether an autonomous agent can be tricked into abusing an external API. And they do not verify that your data is safe from silent training leaks.&lt;/p&gt;

&lt;p&gt;Frameworks like the &lt;strong&gt;NIST AI Risk Management Framework (RMF)&lt;/strong&gt; emphasize four functions: &lt;em&gt;Govern, Map, Measure, and Manage&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;To turn those guidelines into production reality, we built an end-to-end engineering lifecycle: &lt;strong&gt;Select, Test, Integrate, and Operate&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here is how the lifecycle works, how it performed in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Four-Stage Engineering Lifecycle
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+--------------------------------------------------------------------------+
|                       The Responsible AI Lifecycle                       |
|                                                                          |
|  1. SELECT     ──▶ Model due diligence, use-case scoping, legal checks   |
|  2. TEST       ──▶ Quantitative benchmarking &amp;amp; agentic red teaming       |
|  3. INTEGRATE  ──▶ Layered guardrails, tool mediation, and HITL gates    |
|  4. OPERATE    ──▶ Real-time telemetry, user feedback &amp;amp; regression tests |
+--------------------------------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h3&gt;
  
  
  Stage 1: SELECT (Model Vetting &amp;amp; Risk Scoping)
&lt;/h3&gt;

&lt;p&gt;Before any foundation model is approved for platform use, it undergoes structured due diligence:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Use-Case Scoping&lt;/strong&gt;: Document the intended application. Identify potential risks (safety, fairness, security, operational impact) specific to that domain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Documentation Review&lt;/strong&gt;: Systematically review model cards, technical papers, and independent evaluations. Verify whether training datasets, curation processes, and known limitations are clearly disclosed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Contractual &amp;amp; Legal Scrutiny&lt;/strong&gt;: Review provider terms of service. Confirm that prompt inputs and completions are strictly excluded from model training, and verify regional data residency commitments.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Stage 2: TEST (Technical Evaluation &amp;amp; Red Teaming)
&lt;/h3&gt;

&lt;p&gt;Once a model is vetted, it is subjected to objective testing in an isolated environment before production integration:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Harm &amp;amp; Bias Benchmarks&lt;/strong&gt;: Quantitatively evaluate the model against standardized benchmarks to measure toxicity, stereotyping, and fairness across demographic categories.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic Red Teaming&lt;/strong&gt;: Test the model inside an agentic simulation. Security engineers craft malicious goals to see if the agent can be manipulated into chaining tools, escalating privileges, or executing unauthorized actions through indirect prompt injection.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Stage 3: INTEGRATE (Architectural Safeguards &amp;amp; Mediation)
&lt;/h3&gt;

&lt;p&gt;Passing the test stage clears a model for platform integration with defensive engineering:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input and Output Guardrails&lt;/strong&gt;: Deploy a multi-layered shield that inspects prompts, tool payloads, and model completions for jailbreaks, prompt injection, and PII leakage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool Mediation Proxy&lt;/strong&gt;: The language model never connects directly to backend services. A mediation proxy validates caller identity, checks tool allowlists, and enforces parameter schemas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human-in-the-Loop Workflows&lt;/strong&gt;: Configure mandatory human review checkpoints for write-enabled or destructive operations.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Stage 4: OPERATE (Continuous Observation &amp;amp; Governance)
&lt;/h3&gt;

&lt;p&gt;Responsible AI is a continuous operational discipline:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Real-Time Guardrail Monitoring&lt;/strong&gt;: Track filter trigger rates, blocked injection attempts, and latency spikes across live traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;User Feedback Loops&lt;/strong&gt;: Collect explicit ratings (thumbs up/down) and implicit signals (copying text or abandoning threads) to identify model drift or user frustration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automated Regression Testing&lt;/strong&gt;: When new documents are ingested into knowledge libraries or prompt templates are modified, run automated evaluation test sets to verify that accuracy and safety metrics remain above target thresholds.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Closing the Responsibility Gap&lt;/strong&gt;: Proactively testing models during the Select and Test stages prevented vulnerable or ungrounded models from ever reaching production, rather than discovering flaws after launch.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standardized Reusable Foundations&lt;/strong&gt;: Centralizing the lifecycle at the platform layer meant application teams did not have to negotiate vendor contracts, build custom toxicity filters, or invent their own evaluation suites. They inherited an audited, compliant foundation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Targeted Agentic Red Teaming&lt;/strong&gt;: Simulating multi-step tool misuse uncovered edge cases that simple single-prompt filters missed. We identified and patched subtle tool chaining risks long before production release.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data-Driven Release Gates&lt;/strong&gt;: Product teams knew the exact pass thresholds required for publication (such as scoring above 95% faithfulness on golden test sets). This turned safety from an ambiguous debate into a clear engineering milestone.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Vendor Model Point Releases&lt;/strong&gt;: Hyperscalers frequently update model weights or retire older versions with short notice. Even minor updates can alter an agent's reasoning style or break tool calling schemas. Run automated regression test sets on schedules to catch upstream model drift early.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrail Latency Overhead&lt;/strong&gt;: Layering multiple safety checks (content filtering, PII redaction, prompt shields) can add hundreds of milliseconds of latency. Run lightweight checks in parallel and optimize caching so safety does not compromise user responsiveness.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synthetic Test vs. Real-World Prompts&lt;/strong&gt;: Evaluation test sets can easily reflect what engineers expect rather than how actual users speak. Continuously update your evaluation datasets using anonymized, sanitized production queries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Over-Filtering False Positives&lt;/strong&gt;: Safety filters configured with overly sensitive thresholds will block benign technical or medical terminology. Calibrate filter sensitivity against your specific industry vocabulary to avoid frustrating legitimate users.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>security</category>
    </item>
    <item>
      <title>Pioneering ISO 42001: The Engineering Blueprint for an AI Management System</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:17:14 +0000</pubDate>
      <link>https://dev.to/dks/pioneering-iso-42001-the-engineering-blueprint-for-an-ai-management-system-51dd</link>
      <guid>https://dev.to/dks/pioneering-iso-42001-the-engineering-blueprint-for-an-ai-management-system-51dd</guid>
      <description>&lt;p&gt;For decades, enterprise software engineering operated under established security frameworks: ISO/IEC 27001 and SOC 2 Type II.&lt;/p&gt;

&lt;p&gt;Those frameworks assume deterministic systems. If you authenticate a user, check an access rule, and query a database, the code behaves predictably every time. Security audits focused on access controls, data encryption, network firewalls, and backup recovery tests.&lt;/p&gt;

&lt;p&gt;Generative AI and autonomous agents break this deterministic model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Models produce non-deterministic outputs.&lt;/li&gt;
&lt;li&gt;Autonomous agents decide their own multi-step tool execution paths.&lt;/li&gt;
&lt;li&gt;System inputs include unstructured prompts and external retrieved documents.&lt;/li&gt;
&lt;li&gt;Systems rely on third-party cloud foundation models outside corporate network boundaries.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In late 2023, the International Organization for Standardization published &lt;strong&gt;ISO/IEC 42001:2023&lt;/strong&gt;. It is the world's first certifiable standard for an &lt;strong&gt;Artificial Intelligence Management System (AIMS)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When we set out to build an enterprise AI platform, we realized that security compliance alone was not enough. We needed an engineering blueprint to achieve ISO 42001 certification.&lt;/p&gt;

&lt;p&gt;Here is the idea, how the architecture works, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: The Shared Responsibility Governance Model
&lt;/h2&gt;

&lt;p&gt;The biggest mistake organizations make with ISO 42001 is treating it as an annual paper audit owned solely by legal and compliance teams.&lt;/p&gt;

&lt;p&gt;In a fast-moving engineering organization, compliance must be built directly into the platform architecture.&lt;/p&gt;

&lt;p&gt;To solve this, we used a concept that cloud architects already know well: the &lt;strong&gt;Shared Responsibility Model&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The Shared Responsibility Model is not new. Cloud hyperscalers like Amazon Web Services (AWS) and Microsoft Azure popularized it over a decade ago for cloud infrastructure. In the cloud world, the provider takes responsibility for the &lt;strong&gt;"security OF the cloud"&lt;/strong&gt; (the physical facilities, hardware, hypervisors, and core networking). The customer takes responsibility for the &lt;strong&gt;"security IN the cloud"&lt;/strong&gt; (their guest operating systems, identity permissions, network traffic rules, and application data).&lt;/p&gt;

&lt;p&gt;We applied that exact, proven architectural division to enterprise generative AI:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Security and Governance OF the Platform (Platform Team Responsibility)&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;Providing secure, isolated multi-tenant compute and networking.&lt;/li&gt;
&lt;li&gt;Enforcing zero-data-retention agreements and data residency across foundation model providers.&lt;/li&gt;
&lt;li&gt;Capturing immutable event logs of all agent runs, tool calls, and model guardrail interventions.&lt;/li&gt;
&lt;li&gt;Supplying automated evaluation test runners and synthetic non-production data sandboxes.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security and Governance IN the Platform (Application Team Responsibility)&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;Defining the specific intended use, target users, and operational boundaries of each agent.&lt;/li&gt;
&lt;li&gt;Conducting an AI System Impact Assessment before deploying to production.&lt;/li&gt;
&lt;li&gt;Curating, classifying, and verifying the quality of documents placed into knowledge libraries.&lt;/li&gt;
&lt;li&gt;Selecting appropriate tool allowlists and setting human-in-the-loop review gates.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This division prevents bottlenecks. Just like developers do not worry about physical data center locks when deploying an app to AWS or Azure, our application teams do not worry about model provider data retention, trace encryption, or network egress isolation. The platform gives engineering teams pre-certified guardrails by default, while holding application creators accountable for what they inject into their agents.&lt;/p&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Unlocking Enterprise Customer Trust&lt;/strong&gt;: Many enterprise customers hesitate to adopt generative AI out of fear of hallucinations, intellectual property leakage, or bias. Becoming an early adopter of ISO 42001 provided customers and third-party auditors with verifiable evidence of responsible AI engineering.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standardized Risk Vocabulary&lt;/strong&gt;: Instead of subjective debates about whether an AI feature was "safe to ship," teams used standardized assessment criteria. Discussions shifted from opinions to measurable risk levels, mitigation plans, and residual risk sign-offs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guardrails as Architecture, Not Checklists&lt;/strong&gt;: By baking logging, data sanitization, and model gateway protections directly into shared infrastructure, individual developer teams inherited ISO 42001 controls automatically without writing boilerplate compliance logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decoupled Velocity&lt;/strong&gt;: Application teams could launch new AI agents in days because the underlying compute, storage, and vendor governance controls were already audited and approved at the platform layer.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Compliance Theater and Bureaucracy&lt;/strong&gt;: If building an internal proof-of-concept requires weeks of governance paperwork, developers will bypass the platform and build shadow AI. Establish tiered risk pathways: low-risk internal advisory experiments should have lightweight, automated checks, reserving formal multi-party audits for high-impact or customer-facing agents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scope Creep Across Non-AI Systems&lt;/strong&gt;: ISO 42001 is designed specifically for artificial intelligence components. Keep a clear boundary between your AI Management System and your existing ISO 27001 information security controls so audits remain focused on AI-specific risks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Continuous Living Standard&lt;/strong&gt;: Unlike SOC 2, which often focuses on historical look-back periods, an AI management system requires continuous monitoring. As models evolve, prompts change, and new tools are connected, your governance mechanisms must adapt in real time.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>architecture</category>
      <category>security</category>
      <category>softwareengineering</category>
    </item>
    <item>
      <title>The Token Vault: Secure OAuth Delegation and Identity for AI Agents</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:14:20 +0000</pubDate>
      <link>https://dev.to/dks/the-token-vault-secure-oauth-delegation-and-identity-for-ai-agents-m68</link>
      <guid>https://dev.to/dks/the-token-vault-secure-oauth-delegation-and-identity-for-ai-agents-m68</guid>
      <description>&lt;p&gt;When building an enterprise AI agent, you will quickly hit a critical security dilemma:&lt;/p&gt;

&lt;p&gt;A user asks an agent to inspect a document in their cloud storage drive and update an issue in their project management tracker. To perform these tasks, the agent must call external APIs on behalf of that specific user.&lt;/p&gt;

&lt;p&gt;How do you give the agent the necessary authorization tokens to execute these actions?&lt;/p&gt;

&lt;p&gt;In early prototypes, developers often inject the user's raw Bearer token directly into the LLM system prompt or into the tool arguments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Putting raw credentials into an LLM prompt or tool argument is an unacceptable security hazard.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If an attacker executes a prompt injection attack (such as hiding instructions inside a shared document that direct the model to output its system prompt), the model can leak the user's secret keys into the chat window. Furthermore, prompts and completions are logged across monitoring tools and vector caches, exposing sensitive credentials in plain text.&lt;/p&gt;

&lt;p&gt;To solve this, we architected the &lt;strong&gt;Token Vault&lt;/strong&gt;. It enables secure On-Behalf-Of (OBO) delegation so agents can run tools without ever seeing the user's raw secrets.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: Opaque Identity References
&lt;/h2&gt;

&lt;p&gt;The core idea of the Token Vault is to decouple the &lt;strong&gt;agent's reasoning engine&lt;/strong&gt; from the &lt;strong&gt;credential layer&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Instead of giving credentials to the language model, all secrets live in an isolated, encrypted vault service:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;User Consent &amp;amp; Vault Storage&lt;/strong&gt;: When a user authorizes an integration, standard OAuth authorization code flows occur between the user's browser and the Token Vault. The vault encrypts and stores the access and refresh tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Opaque Token References&lt;/strong&gt;: The vault generates a randomized, opaque reference identifier (a non-sensitive handle) representing that active connection.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt-Free Execution&lt;/strong&gt;: The agent only ever sees and passes the opaque handle. When the agent calls a tool, our backend tool broker intercepts the call, validates the user's session, exchanges the opaque handle for a short-lived token inside the secure network perimeter, executes the API call, and returns the result.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The language model, the prompt context, and the observability logs never touch a real secret key.&lt;/p&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Zero Credential Leakage via Prompt Injections&lt;/strong&gt;: Even if an attacker completely compromises an agent's reasoning chain through indirect prompt injection, the model cannot leak the token because it does not have it. The model only knows an opaque reference handle that is useless outside of the authenticated session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automated Refresh Management&lt;/strong&gt;: Third-party access tokens typically expire in one hour. Because the Token Vault manages the OAuth lifecycle, it automatically uses the stored refresh token to obtain a fresh access token without interrupting the user's conversation or requiring re-authentication.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Strict Caller-Identity Binding&lt;/strong&gt;: The Token Vault enforces cryptographic ownership checks. If User B inspects an agent's session and attempts to invoke a tool using User A's reference handle, the vault rejects the exchange with a permission error.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clean Observability Compliance&lt;/strong&gt;: OpenTelemetry traces and audit logs capture full tool execution details without triggering data-loss-prevention alarms or exposing user credentials in telemetry dashboards.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Granular Scope Containment&lt;/strong&gt;: When users connect external accounts, request the minimum necessary OAuth scopes. If an agent only needs to read documents, never request write or admin permissions. If an agent needs write access, implement step-up confirmation so users explicitly approve the elevated scope.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session and Token Revocation&lt;/strong&gt;: If a user logs out of your enterprise platform or disconnects a third-party service, your Token Vault must immediately invalidate the cached credentials and revoke the upstream OAuth grant. Disconnecting one tool should never leave orphaned active tokens in storage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rate Limits on Refresh Flows&lt;/strong&gt;: Third-party identity providers place rate limits on token refresh endpoints. If hundreds of concurrent agent runs attempt to refresh tokens simultaneously for the same user, you can trigger provider throttling. Implement distributed caching and locking around refresh token exchanges.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Handling Disconnected State Gracefully&lt;/strong&gt;: When an upstream token expires and cannot be refreshed (for example, if the user changed their enterprise password), the tool broker must return a clean, structured re-authentication prompt to the chat client rather than crashing the agent with an unhandled exception.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>llm</category>
      <category>security</category>
    </item>
    <item>
      <title>Scaling 10,000+ AI Batch Jobs with EKS Auto Mode and Karpenter</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:14:18 +0000</pubDate>
      <link>https://dev.to/dks/scaling-10000-ai-batch-jobs-with-eks-auto-mode-and-karpenter-2l6l</link>
      <guid>https://dev.to/dks/scaling-10000-ai-batch-jobs-with-eks-auto-mode-and-karpenter-2l6l</guid>
      <description>&lt;p&gt;Building a Retrieval-Augmented Generation (RAG) platform looks deceptively simple in proof-of-concept tutorials. You read a PDF, split it into chunks, call an embedding API, and store the vectors in a database.&lt;/p&gt;

&lt;p&gt;In an enterprise, scale changes the equation completely.&lt;/p&gt;

&lt;p&gt;When project teams upload hundreds of thousands of technical documents, contracts, and manuals, your platform experiences massive, unpredictable traffic spikes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;During peak business hours, users queue 20,000 documents for indexing.&lt;/li&gt;
&lt;li&gt;Overnight and on weekends, the queue sits empty.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you provision enough on-demand cloud servers to process the peak 20,000-document backlog immediately, you pay thousands of dollars for idle compute during off-peak periods.&lt;/p&gt;

&lt;p&gt;If you downsize your cluster to save money, documents sit in queue backlogs for hours, frustrating users who expect instant search availability.&lt;/p&gt;

&lt;p&gt;Here is the idea we used to solve this on Amazon EKS, how it performed in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: Event-Driven Queue Scaling with Spot Compute
&lt;/h2&gt;

&lt;p&gt;To achieve both high throughput and extreme cost efficiency, we combined three architectural patterns:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Decoupled Asynchronous Ingestion&lt;/strong&gt;: Document uploads are decoupled from processing. When files arrive, lightweight ingest endpoints place processing jobs onto high-throughput cloud queues.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backlog-Driven Pod Autoscaling (KEDA)&lt;/strong&gt;: Rather than scaling based on lagging CPU or memory metrics, we scale worker pods based directly on queue depth. If thousands of jobs land in the queue, autoscalers immediately request hundreds of worker pods within seconds, and scale completely down to zero when the queue clears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Just-In-Time Node Provisioning (Karpenter Spot)&lt;/strong&gt;: When hundreds of worker pods become pending, an intelligent node provisioner communicates directly with the cloud fleet API. It dynamically selects from a diverse pool of discounted Spot instance types and architectures, launching right-sized nodes in under 45 seconds.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;70% to 90% Cost Reduction&lt;/strong&gt;: Spot instances trade availability guarantees for massive discounts. Because batch document parsing is naturally fault-tolerant, running batch workers on Spot instances reduced our compute bill by up to 90% compared to on-demand instances.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instant Queue Draining&lt;/strong&gt;: When thousands of documents are uploaded simultaneously, the autoscaler skips gradual ramp-up intervals, jumping from zero to hundreds of concurrent pods. The node provisioner launches dozens of varied EC2 instances in parallel, draining massive backlogs in minutes instead of hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;True Scale-to-Zero Efficiency&lt;/strong&gt;: When queues empty, pods scale to zero, and the node provisioner automatically consolidates and terminates the underlying worker nodes within seconds. We pay zero compute costs when no documents are being processed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Architectural Flexibility Across Chipsets&lt;/strong&gt;: By allowing the node provisioner to schedule across both standard x86 and ARM-based Graviton chipsets, the system automatically selects the cheapest available capacity in each cloud region.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Handling Spot Interruptions Gracefully&lt;/strong&gt;: Spot instances can be reclaimed by cloud providers with a two-minute warning. If a node is reclaimed while a worker is halfway through parsing a 100-page document, the job must not be lost. Configure queue visibility timeouts and dead-letter queues so that interrupted jobs become visible for other workers to retry automatically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Avoid Throttling Downstream Embedding APIs&lt;/strong&gt;: When scaling from 2 pods to 500 pods in 60 seconds, your workers will make hundreds of concurrent calls to your embedding model endpoints. If you do not configure rate-limiting and connection pooling, you will overwhelm your upstream model gateway with HTTP 429 errors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Spot Pool Diversification&lt;/strong&gt;: If you restrict your node provisioner to only one or two specific instance types (e.g. only &lt;code&gt;c6i.4xlarge&lt;/code&gt;), your cluster will experience provisioning delays when cloud capacity pools run low. Always configure your provisioners to draw from dozens of instance families across multiple availability zones.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node Startup Optimization&lt;/strong&gt;: A pod cannot start until its container image is pulled. For large AI worker containers (which often bundle OCR libraries and Python dependencies), pull times can take several minutes. Optimize your container images by stripping unnecessary packages, using lightweight base images, and taking advantage of node caching or fast-launch disk snapshots.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>aws</category>
      <category>kubernetes</category>
      <category>rag</category>
    </item>
    <item>
      <title>Pioneering ISO 42001: Event Logging and Reasoning Trace Auditability</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:12:11 +0000</pubDate>
      <link>https://dev.to/dks/pioneering-iso-42001-event-logging-and-reasoning-trace-auditability-2nl4</link>
      <guid>https://dev.to/dks/pioneering-iso-42001-event-logging-and-reasoning-trace-auditability-2nl4</guid>
      <description>&lt;p&gt;In traditional web applications, auditing is simple. You record who logged in, which database records were created or modified, and which API endpoints returned errors.&lt;/p&gt;

&lt;p&gt;In an autonomous agent platform, auditing is much more demanding.&lt;/p&gt;

&lt;p&gt;When an AI agent runs, it does not execute a hardcoded path. It receives a prompt, considers its available tools, invokes one or more external APIs, inspects the responses, synthesizes an answer, and presents it to the user.&lt;/p&gt;

&lt;p&gt;If that agent makes an erroneous decision or executes an unauthorized tool action, you cannot simply look at an HTTP 500 status code. You must be able to reconstruct:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Exactly what was the user's prompt?&lt;/li&gt;
&lt;li&gt;Which system prompt and agent configuration version was in effect?&lt;/li&gt;
&lt;li&gt;Which foundation model generated the thought?&lt;/li&gt;
&lt;li&gt;What tools were called, with what arguments, and what data did the tools return?&lt;/li&gt;
&lt;li&gt;Did safety guardrails intervene or flag any tokens?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This level of auditability is mandated by &lt;strong&gt;ISO/IEC 42001 Control A.6.2.8 (AI System Recording of Event Logs)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Here is how we architected event logging and reasoning trace auditability, how it worked in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: Life-Cycle Logging and Immutable Configurations
&lt;/h2&gt;

&lt;p&gt;Under Control A.6.2.8, organizations must make a recorded decision on which life-cycle phases have event logging enabled, how long records are kept, and how they can be retrieved during an audit.&lt;/p&gt;

&lt;p&gt;We structured our logging architecture around three core concepts:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+--------------------------------------------------------------------------+
|                  The Complete Agent Trace Lifecycle                      |
|                                                                          |
|  [ User Prompt ]                                                         |
|         |                                                                |
|         v                                                                |
|  [ Immutable Agent Version ] ──▶ (System prompt, allowed tools, models)   |
|         |                                                                |
|         v                                                                |
|  [ Edge Safety Guardrail ]   ──▶ (Logged: Injection &amp;amp; content screening) |
|         |                                                                |
|         v                                                                |
|  [ Tool Invocation Steps ]   ──▶ (Logged: Arguments, durations, results)  |
|         |                                                                |
|         v                                                                |
|  [ Final Response &amp;amp; Tokens ] ──▶ (Logged: TTFT, TPOT, token costs)       |
+--------------------------------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1. Immutable Configuration Versions as Change Records
&lt;/h3&gt;

&lt;p&gt;Agent behavior is governed by prompts, model choices, tool allowlists, and execution parameters. In our architecture, these configurations are strictly immutable.&lt;/p&gt;

&lt;p&gt;When an engineer or creator edits an agent, the system never overwrites the existing configuration in place. It creates a new, numbered version in an unbroken chain.&lt;/p&gt;

&lt;p&gt;Because every single execution records the exact configuration version that served it, any production output can be traced back to the exact system prompt in force at that second, and that prompt back to the engineer who saved it.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Multi-Step Execution Tracing
&lt;/h3&gt;

&lt;p&gt;Every agent interaction records the complete ordered exchange:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The initial user input and session context.&lt;/li&gt;
&lt;li&gt;Intermediate reasoning thoughts and subagent handoffs.&lt;/li&gt;
&lt;li&gt;Exact tool calls, payload arguments, and returned tool data.&lt;/li&gt;
&lt;li&gt;Prompt guardrail screening results (such as whether content filters intervened).&lt;/li&gt;
&lt;li&gt;Detailed token consumption and latency metrics.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Edge PII Sanitization
&lt;/h3&gt;

&lt;p&gt;Logging complete reasoning chains introduces a severe privacy risk: what if a user pastes a social security number, API key, or personal contact info into the conversation?&lt;/p&gt;

&lt;p&gt;Our telemetry collectors run automated sanitization filters at the edge. Sensitive entity patterns are replaced with redaction tokens before the event is permanently written to our audit data warehouse.&lt;/p&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Instant Root-Cause Investigation&lt;/strong&gt;: When a user reported an unexpected response, support and engineering teams did not have to guess what happened. Opening the trace view showed the complete step-by-step reasoning trajectory in seconds.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Demonstrable Audit Compliance&lt;/strong&gt;: When external auditors asked to see evidence of operational monitoring, we produced complete, queryable logs showing the lifecycle of agents from testing through production usage and eventual decommissioning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety Guardrail Visibility&lt;/strong&gt;: Logging safety filter outcomes gave our security team clear visibility into malicious prompt injection attempts and jailbreak patterns across our multi-tenant platform.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliable Rollbacks&lt;/strong&gt;: Because configuration versions are immutable, reverting an agent to a known good state after an issue is instantaneous. Operators simply point the active alias to the prior configuration version.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Log Volume and Storage Costs&lt;/strong&gt;: Logging full prompt and response payloads for millions of daily queries produces massive data volumes. Separate real-time operational APM (which only needs metrics and status codes) from deep audit logging, storing full traces in cost-effective columnar storage with defined retention policies.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Over-Logging in Development&lt;/strong&gt;: Emitting high-volume verbose tracing during local development or unit testing can clutter telemetry pipelines and incur unnecessary cloud costs. Define clear logging rules per life-cycle phase so that ephemeral tests do not flood production audit stores.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Redaction Latency&lt;/strong&gt;: Real-time PII regex and entity masking must be lightweight. If your sanitization engine introduces hundreds of milliseconds of overhead to every streaming chunk, user responsiveness will suffer.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Handling Deleted Resources&lt;/strong&gt;: If a user deletes an agent or a document library, your audit trail must retain the deletion event and historical run records even after the active resource is gone. Never hard-delete the audit logs associated with removed services.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>architecture</category>
      <category>security</category>
    </item>
    <item>
      <title>Pioneering ISO 42001: Data Quality and Provenance in RAG Systems</title>
      <dc:creator>Dinesh Kumar Sarangapani</dc:creator>
      <pubDate>Sat, 10 Oct 2026 14:12:08 +0000</pubDate>
      <link>https://dev.to/dks/pioneering-iso-42001-data-quality-and-provenance-in-rag-systems-2ni2</link>
      <guid>https://dev.to/dks/pioneering-iso-42001-data-quality-and-provenance-in-rag-systems-2ni2</guid>
      <description>&lt;p&gt;In traditional software, data quality is governed by relational database schemas, unique constraints, and foreign keys. If a record has the wrong data type, the database rejects the write.&lt;/p&gt;

&lt;p&gt;In Retrieval-Augmented Generation (RAG) and agentic AI, data quality is far more complex:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Documents arrive as messy PDFs, scanned images, Word files, and raw markdown.&lt;/li&gt;
&lt;li&gt;Parsers extract text and tables with varying degrees of accuracy.&lt;/li&gt;
&lt;li&gt;Text chunking algorithms split sentences across boundaries, occasionally destroying semantic meaning.&lt;/li&gt;
&lt;li&gt;Stale or outdated documents remain indexed alongside current versions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your vector database contains incomplete, corrupted, or unvetted text, your language model will generate flawed answers no matter how advanced its reasoning capabilities are.&lt;/p&gt;

&lt;p&gt;ISO/IEC 42001 explicitly addresses this challenge in &lt;strong&gt;Annex A.7 (Data for AI Systems)&lt;/strong&gt;. It requires organizations to establish rigorous controls around data quality, data preparation, sourcing, and provenance.&lt;/p&gt;

&lt;p&gt;Here is how we operationalized data quality and provenance for our enterprise AI platform, how it worked in production, and what to watch out for.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Idea: Measurable Data Quality Dimensions
&lt;/h2&gt;

&lt;p&gt;In the cloud world, providers like AWS and Azure made the Shared Responsibility Model standard for infrastructure. We applied that exact concept to RAG data quality:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Quality OF the Platform&lt;/strong&gt;: The platform engineering team ensures file format validation, high parsing fidelity, chunk tracking without silent truncation, and cryptographic provenance tagging.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality IN the Platform&lt;/strong&gt;: The agent creator or content author ensures that the uploaded source documents are factually accurate, current, and classified properly according to data policies.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Under this model, we defined four measurable data quality dimensions for all text and documents ingested into the platform:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;+--------------------------------------------------------------------------+
|                      RAG Data Quality Dimensions                         |
|                                                                          |
|  1. Accuracy (Parsing)  ──▶ Was text extracted without corruption?       |
|  2. Completeness        ──▶ Were all valid sections chunked and indexed? |
|  3. Consistency         ──▶ Do files conform to supported schemas?       |
|  4. Timeliness          ──▶ Is data tagged with clear freshness dates?   |
+--------------------------------------------------------------------------+
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1. Ingestion Hygiene and Validation
&lt;/h3&gt;

&lt;p&gt;Before a document enters our vector pipeline, it passes through structural validation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Supported format verification (rejecting malformed or corrupted file structures).&lt;/li&gt;
&lt;li&gt;Security screening for malicious macro code and embedded payloads.&lt;/li&gt;
&lt;li&gt;Extraction quality checks (flagging unsearchable scanned PDFs that require OCR).&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Data Provenance and Lineage
&lt;/h3&gt;

&lt;p&gt;Every chunk of text stored in our vector database retains full provenance metadata:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Original source file name and content hash.&lt;/li&gt;
&lt;li&gt;Author, ingestion timestamp, and last-modified date.&lt;/li&gt;
&lt;li&gt;Explicit permissions and access control tags.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When an AI agent generates an answer, it cites the exact document chunk and provenance reference, allowing users to verify the source material directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Synthetic Non-Production Isolation
&lt;/h3&gt;

&lt;p&gt;A strict control in our compliance framework is the &lt;strong&gt;prohibition of production customer data in non-production environments&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;When testing new retrieval algorithms, training evaluation benchmarks, or building demo agents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Development and staging environments use synthetic or public datasets only.&lt;/li&gt;
&lt;li&gt;Ingestion pipelines are isolated by environment so that production document stores cannot be read by test runners.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  How It Worked Well
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Eliminating Garbage-In Hallucinations&lt;/strong&gt;: Enforcing strict parsing validation prevented corrupted text (such as garbled table formatting or broken OCR text) from polluting vector indexes. High-quality parsing led to a measurable improvement in RAG retrieval accuracy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Transparent Source Verification&lt;/strong&gt;: Because every chunk carries cryptographic provenance metadata, users and auditors can trace every generated paragraph back to its exact page and paragraph in the source document.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Audit-Ready Data Registers&lt;/strong&gt;: When compliance auditors asked how we managed intellectual property and data licensing, our automated data registers provided immediate reports on data sources, retention schedules, and access boundaries.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Clean Stale-Document Pruning&lt;/strong&gt;: Storing last-modified metadata allowed platform teams to identify and prune outdated policy documents, ensuring agents did not cite obsolete operating procedures.&lt;/li&gt;
&lt;/ol&gt;




&lt;h2&gt;
  
  
  What to Watch Out For
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;PDF Parsing Traps&lt;/strong&gt;: Complex multi-column PDFs and dense tables are notorious for confusing standard text extractors. If columns are merged horizontally, the extracted text becomes nonsense. Use specialized layout-aware document parsers and test them against your organization's actual document templates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Silent Ingestion Truncation&lt;/strong&gt;: Some document processors quietly truncate files if they exceed token or memory limits, leaving the second half of a 100-page document unindexed. Always log total extracted character counts and verify that the number of generated chunks matches the full source volume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Permission Desynchronization&lt;/strong&gt;: If a user updates access permissions on a source folder in an enterprise cloud drive, does your vector store know? If permissions do not sync dynamically, an agent might retrieve sensitive documents and expose them to unauthorized users. Ensure retrieval queries filter chunks against real-time user access control lists.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synthetic Data Drift&lt;/strong&gt;: While synthetic data is essential for protecting customer privacy in test environments, synthetic datasets can easily drift from the messy realities of production documents. Regularly calibrate your synthetic generation pipelines to reflect production formatting and terminology.&lt;/li&gt;
&lt;/ol&gt;

</description>
      <category>ai</category>
      <category>database</category>
      <category>llm</category>
      <category>rag</category>
    </item>
  </channel>
</rss>
