<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rithika</title>
    <description>The latest articles on DEV Community by Rithika (@rithika_7575).</description>
    <link>https://dev.to/rithika_7575</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4153637%2Fe6b40232-f40a-4801-929c-d6a98526c605.png</url>
      <title>DEV Community: Rithika</title>
      <link>https://dev.to/rithika_7575</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/rithika_7575"/>
    <language>en</language>
    <item>
      <title>Weekend Challenge: Model Substitution Governance &amp; Audit Platform for Dynamic LLM Gateways</title>
      <dc:creator>Rithika</dc:creator>
      <pubDate>Mon, 05 Oct 2026 05:15:06 +0000</pubDate>
      <link>https://dev.to/rithika_7575/weekend-challenge-model-substitution-governance-audit-platform-for-dynamic-llm-gateways-17be</link>
      <guid>https://dev.to/rithika_7575/weekend-challenge-model-substitution-governance-audit-platform-for-dynamic-llm-gateways-17be</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for the &lt;a href="https://dev.to/challenges/hacktoberfest-weekend-2026-10-01"&gt;Hacktoberfest Weekend Challenge: Build for a Friend&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Built
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;I built the Model Substitution Governance &amp;amp; Audit Platform for my close friend, an AI/MLOps engineer who manages multi-agent systems and dynamic LLM routing gateways in production. Like many teams running autonomous agent pipelines, their infrastructure relies on dynamic gateways (like LiteLLM, vLLM, or custom fallback routers) to swap models on the fly whenever high-traffic spikes, token budgets, or provider rate limits hit.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The hidden pain point? Silent model substitutions. When a mission-critical agent requests a model with a massive context window and complex reasoning capabilities (such as Claude 3.5 Sonnet or Llama 3 70B), the gateway might silently downgrade the call to a smaller or cheaper fallback (like Mistral 7B or GPT-4o Mini) to keep latency low. This causes silent context window truncation, hallucinated outputs, loss of long-document context, and unexpected compliance violations when prompts get routed to unapproved external providers. My friend had zero unified audit trail or alerting system to know when, why, or how often these downgrades were degrading their agent's downstream outputs.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;To solve this, I designed an end-to-end governance and auditing platform. It pairs a zero-latency, non-blocking Python Interceptor SDK (governance-interceptor) with a high-throughput FastAPI Cloud Tracker, an automated Capability Risk Assessor Engine, and a real-time Vercel Glassmorphism Dashboard. &lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;The system automatically computes context downgrade percentages, assigns risk severity ratings (LOW, MEDIUM, HIGH, CRITICAL), verifies agent provider whitelists, and produces retroactive compliance audit reports.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Demo
&lt;/h2&gt;

&lt;p&gt;Experience the full live platform and test simulated gateway routing in real time:&lt;br&gt;
Live Web Dashboard:model-substitution-governance-event.vercel.app&lt;br&gt;
&lt;a href="https://model-substitution-governance-event.vercel.app" rel="noopener noreferrer"&gt;https://model-substitution-governance-event.vercel.app&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Cloud API Engine &amp;amp; Swagger Specs:model-substitution-governance-event.onrender.com/docs&lt;br&gt;
&lt;a href="https://model-substitution-governance-event.onrender.com/docs" rel="noopener noreferrer"&gt;https://model-substitution-governance-event.onrender.com/docs&lt;/a&gt;&lt;br&gt;
SDK Release Package: v1.0.0 Release Archive&lt;br&gt;
&lt;a href="https://github.com/Rithika-Gurusamy/Model-Substitution-Governance-Event-Ps---8.2-/releases/download/v1.0.0/governance-interceptor-v1.0.0.zip" rel="noopener noreferrer"&gt;https://github.com/Rithika-Gurusamy/Model-Substitution-Governance-Event-Ps---8.2-/releases/download/v1.0.0/governance-interceptor-v1.0.0.zip&lt;/a&gt;&lt;br&gt;
 --&lt;/p&gt;

&lt;p&gt;The dashboard includes a built-in "Try Demo" Interactive Simulator modal that lets anyone trigger mock gateway substitutions with one click to observe real-time risk scoring, context delta computations, and whitelist policy alerts without writing a line of code.&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;The complete source code for the backend, frontend dashboard, and interceptor SDK is open source on GitHub:&lt;br&gt;
&lt;/p&gt;
&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/Rithika-Gurusamy" rel="noopener noreferrer"&gt;
        Rithika-Gurusamy
      &lt;/a&gt; / &lt;a href="https://github.com/Rithika-Gurusamy/Model-Substitutions-Governance-Platform" rel="noopener noreferrer"&gt;
        Model-Substitutions-Governance-Platform
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      A platform to govern model substituitions in your applications 
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;div class="markdown-heading"&gt;
&lt;h1 class="heading-element"&gt;MODEL SUBSTITUTION GOVERNANCE &amp;amp; AUDIT PLATFORM&lt;/h1&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;Real-Time Monitoring, Capability Risk Assessment, and Compliance Auditing for Dynamic LLM Gateway Model Routing&lt;/strong&gt;&lt;/p&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;LIVE PRODUCTION LINKS&lt;/h2&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Repository&lt;/strong&gt;: &lt;a href="https://github.com/Rithika-Gurusamy/Model-Substitution-Governance-Event-Ps---8.2-" rel="noopener noreferrer"&gt;Rithika-Gurusamy/Model-Substitution-Governance-Event&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Live Web Dashboard&lt;/strong&gt;: &lt;a href="https://model-substitution-governance-event.vercel.app" rel="nofollow noopener noreferrer"&gt;model-substitution-governance-event.vercel.app&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud API Engine&lt;/strong&gt;: &lt;a href="https://model-substitution-governance-event.onrender.com" rel="nofollow noopener noreferrer"&gt;model-substitution-governance-event.onrender.com&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAPI / Swagger Specs&lt;/strong&gt;: &lt;a href="https://model-substitution-governance-event.onrender.com/docs" rel="nofollow noopener noreferrer"&gt;model-substitution-governance-event.onrender.com/docs&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SDK GitHub Release&lt;/strong&gt;: &lt;a href="https://github.com/Rithika-Gurusamy/Model-Substitution-Governance-Event-Ps---8.2-/releases/download/v1.0.0/governance-interceptor-v1.0.0.zip" rel="noopener noreferrer"&gt;v1.0.0 Release Package&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;EXECUTIVE SUMMARY&lt;/h2&gt;
&lt;/div&gt;
&lt;p&gt;Modern AI applications use &lt;strong&gt;LLM Gateways&lt;/strong&gt; (such as LiteLLM, Portkey, or custom routing services) to dynamically route prompt requests based on cost, latency, or rate limits. When a high-capability model (e.g., &lt;code&gt;GPT-4o&lt;/code&gt; or &lt;code&gt;Claude 3.5 Sonnet&lt;/code&gt;) is swapped for a smaller model (e.g., &lt;code&gt;Gemini 1.5 Flash&lt;/code&gt; or &lt;code&gt;GPT-4o Mini&lt;/code&gt;), &lt;strong&gt;silent model substitutions&lt;/strong&gt; occur.&lt;/p&gt;
&lt;p&gt;Without governance tracking:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context Degradation&lt;/strong&gt;: Shrinking context windows (e.g., 200k tokens down to 128k) cause subtle reasoning failures or truncation in multi-turn workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compliance &amp;amp; Policy Violations&lt;/strong&gt;: AI agents may route prompts to unapproved or non-whitelisted model providers in regulated environments…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/Rithika-Gurusamy/Model-Substitutions-Governance-Platform" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;Connecting any LLM gateway or agent pipeline requires just 3 lines of code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;governance_interceptor&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;GovernanceInterceptor&lt;/span&gt;

&lt;span class="n"&gt;interceptor&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;GovernanceInterceptor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;tracker_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://model-substitution-governance-event.onrender.com&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;usr_live_your_key_here&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;Intercept&lt;/span&gt; &lt;span class="n"&gt;gateway&lt;/span&gt; &lt;span class="n"&gt;decisions&lt;/span&gt; &lt;span class="n"&gt;asynchronously&lt;/span&gt; &lt;span class="n"&gt;without&lt;/span&gt; &lt;span class="n"&gt;adding&lt;/span&gt; &lt;span class="n"&gt;streaming&lt;/span&gt; &lt;span class="n"&gt;latency&lt;/span&gt;

&lt;span class="n"&gt;interceptor&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;intercept&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;requested_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Llama-3-70B-Instruct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;actual_model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Mistral-7B-Instruct&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;reason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cost_budget_exceeded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;agent_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Compliance-Audit-Agent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;session_id&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;session-tx-4091&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How I Built It
&lt;/h2&gt;

&lt;p&gt;The platform was architected from the ground up to operate seamlessly with both open-weight models (Llama 3, Mistral, Gemma 2, Qwen) and proprietary model endpoints managed through open-source routing frameworks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Lightweight Python Interceptor SDK (governance-interceptor):&lt;/strong&gt; Designed as an in-memory middleware that hooks into gateway routing decisions. It compares requested_model against actual_model and fires non-blocking asynchronous HTTP background tasks so that LLM response streaming and user token generation speeds remain completely unaffected (zero added latency).&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Capability Risk Engine (FastAPI &amp;amp; Pydantic):&lt;/strong&gt; When an event is ingested, the engine evaluates the metadata of the requested and substituted models (context window sizes, token limits, and model tiers). If a model swap results in a dramatic reduction in context capacity (e.g., shrinking from 128k tokens to 8k tokens), the system quantifies the capability gap and tags the event with a risk rating.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Agent Whitelist &amp;amp; Compliance Service:&lt;/strong&gt; Teams can register AI agents with strict allowed-model whitelists. If a cost-cutting fallback router routes a sensitive agent to a non-whitelisted provider, the system immediately flags the substitution as an unapproved policy breach.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Data Persistence &amp;amp; Multi-Tenant Security:&lt;/strong&gt; Built on PostgreSQL (Supabase) with SQLAlchemy ORM schemas, supporting multi-tenant organization isolation through hashed developer API keys (usr_live_...) and Supabase JWT authentication.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Modern Glassmorphism UI:&lt;/strong&gt;Built with vanilla semantic HTML, modern responsive CSS, and JavaScript. It provides real-time event streaming, KPI metric cards, capability risk distribution charts, and retroactive batch audit generators.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Does Open Innovation Matter?
&lt;/h2&gt;

&lt;p&gt;Why does open innovation matter for what you built?  What did it make possible that a closed API wouldn't? &lt;/p&gt;

&lt;h2&gt;
  
  
  My Agent Session
&lt;/h2&gt;

&lt;p&gt;The system design and iterative refactoring of the multi-tenant auth and compliance audit engine were pair-programmed with AI coding assistance. You can review the repository commits and architecture evolution directly on GitHub: Commit History &amp;amp; Architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prize Categories
&lt;/h2&gt;

&lt;p&gt;Build for a Friend (Built specifically for an AI/MLOps engineer friend running multi-agent LLM gateway infrastructure)&lt;br&gt;
Open Source AI / Open Weight AI Governance&lt;/p&gt;

&lt;p&gt;Submissions: DEV username: &lt;a class="mentioned-user" href="https://dev.to/rithika_7575"&gt;@rithika_7575&lt;/a&gt; &lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>weekendchallenge</category>
      <category>hf26challenge</category>
    </item>
  </channel>
</rss>
